# Small benchmark

**URL:** <https://discourse.julialang.org/t/small-benchmark/17665>\
**Category:** Performance\
**Tags:** benchmark\
**Created:** [November 18, 2018, 3:23am UTC](https://discourse.julialang.org/t/small-benchmark/17665 "2018-11-18T03:23:08Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 18, 2018, 3:23am UTC](https://discourse.julialang.org/t/small-benchmark/17665/1 "2018-11-18T03:23:08Z")

</div>

I’m new to Julia, and was curious to test its famed performance, so today I made a small benchmark comparing it to C, optimized Python and Scala. Some pretty interesting results, with Julia falling a mere 3% behind the best C implementation of the code. I’m impressed!

My test was just calculating the exponential function using the same method found in glibc (“math.h”). Key to achieving the highest speeds was using @fastmath, but interestingly the C was pretty bad without that, I don’t know what optimizations I can be missing there. I’m still writing a blog about it (usually I write too much) but I would like to share the code and results in the forum right away. Any opinions and comments are appreciated.

 ![juliabench-bars](https://global.discourse-cdn.com/julialang/original/3X/e/3/e3f2c1e76e04f76e67ddedb8899333199758ad5c.png)

> <https://gist.github.com/nlw0/e3c94ffd816b5592736c0e822f0424be>

> <https://gist.github.com/nlw0/7ee419b1e1c0fb11c5b5a1f6e0e35f81>

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 18, 2018, 5:11pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/2 "2018-11-18T17:11:11Z")

</div>

I have already published the results from this experiment in a blog

> **[Meeting Julia, a great new alternative for numerical programming — Part I...](https://medium.com/@nwerneck/meeting-julia-a-great-new-alternative-for-numerical-programming-part-i-benchmarking-c03dd3289493)**
>
> This is a two-part article presenting the Julia language. Because its promise of C-like performance is a big issue surrounding it, we’ll…

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [November 19, 2018, 7:43pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/3 "2018-11-19T19:43:24Z")

</div>

Nice article! Some feedback on your benchmark:

- You’re including a call to the built-in `exp` function in the timing, which presumably can vary a bit in performance between the different languages. Is that intentional? I think it’d be more interesting to just benchmark the `myexp` code.
- You’re dividing by `n` for each iteration, which is quite slow. At least in Julia, you can save some time by instead multiplying by a pre-calculated `1/n`. That could explain some of the anomalies you’re seeing with Scala, since the code looks a bit different there (you have `0.6/n` instead of `i/n`, so perhaps the compiler is clever enough to do this optimization for you?).
- How many times did you run this to produce the benchmark numbers? When I run your script over and over again, timings vary by a few percent between runs. So if you’re interested in performance differences as small as 3 %, you should probably run your experiment many times to ensure that the results are statistically sound. Perhaps you accounted for this already.

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 20, 2018, 8:27pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/4 "2018-11-20T20:27:14Z")

</div>

Thanks for reading! I thinking leaving the system exp shouldn’t matter much because in the worse case all implementations would be using the best, well optimized version available, and in the worst case one of them might _not_, but this should definitely count as a demerit to that language. In the end it makes the benchmark a little more “broad-spectrum”, perhaps. About the division, I would also hope the compiler can pick that up, and if a language requires that care it should also be a demerit. But it would be definitely interesting to test if this was the case in any of them.

The tests were done only with a single run, not too careful. Some of the numbers were quite consistent, though. The only care I took was to run the test starting with larger batches of numbers to be more fair to JIT languages — we`re not interested in things like start-up time, after all. Once I have a better idea of how I might improve this benchmark in other ways I’ll definitely make a more careful measurements, and then we can look closer at that 3% difference.

---

<div class="post-metadata">

**Author:** ![e3c6](https://avatars.discourse-cdn.com/v4/letter/e/e79b87/32.png) [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Post date:** [November 20, 2018, 8:49pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/5 "2018-11-20T20:49:46Z")

</div>

Unoptimized julia (without `@fastmath`) is faster than C with `O2`. If that’s a sound result, it’s pretty impressive.  
**Edit:** Did you try the plain `for` without `@simd`?

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 20, 2018, 9:04pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/6 "2018-11-20T21:04:04Z")

</div>

Even better than -O3! It is pretty much the same as the “-O2 -finline-functions” there, and a similar performance was attained by Numba and Scala… Maybe there’s just some secret option I am missing that might also make -Ofast even faster,who knows, but that’s what I got here… Would be great to hear if anyone can reproduce this result. I also did not try Clang yet.

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 20, 2018, 9:06pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/7 "2018-11-20T21:06:41Z")

</div>

The @simd actualy made practically no difference, I should remove it from that code.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [November 20, 2018, 10:28pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/8 "2018-11-20T22:28:45Z")

</div>

> [@xor0110](#):
>
> I thinking leaving the system exp shouldn’t matter much because in the worse case all implementations would be using the best, well optimized version available, and in the worst case one of them might _not_ , but this should definitely count as a demerit to that language. In the end it makes the benchmark a little more “broad-spectrum”, perhaps.

Sure, but it adds an unknown. For example, in Julia, you enable fast math by a decorator on the specific test method. In C you pass it to the compiler. Does that mean that you’re enabling fast math for the built-in `exp` in the C version, but not the Julia version? If so, it’s not a fair comparison. (And even if not, it still leaves me as a reader wondering.)

> [@xor0110](#):
>
> About the division, I would also hope the compiler can pick that up, and if a language requires that care it should also be a demerit.

But you’ve coded it differently in the different languages. As I wrote, in Julia you have `i/n` and in some other languages you have `0.6/n`. These are two very different expressions; the latter is a constant within the loop and can easily be optimized by a compiler. The former is not. In fact, if I change the Julia implementation to use `0.6/n`, the run time for 1e9 iterations (for your method only, not the built-in `exp`) goes from 10 seconds to 6.7 seconds(!) So this likely explains why it seemed like Scala was faster than the other languages.

> [@xor0110](#):
>
> The tests were done only with a single run, not too careful.

For Julia, take a look at [BenchmarkTools](https://github.com/JuliaCI/BenchmarkTools.jl). Or for something reproducible between various languages, consider running your test 100 times and reporting the minimum and median times.

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 21, 2018, 9:24am UTC](https://discourse.julialang.org/t/small-benchmark/17665/9 "2018-11-21T09:24:15Z")

</div>

Indeed, the Scala loop is not consistent with the others, I’ll fix that in the next iteration, thanks!

I use functions like exp a lot in my code, so I’m definitely interesting in knowing how fast is my code with that, regardless of the reasons. It would be certainly interesting to know all these details for sure, but we would need multiple specific tests to understand that and I only had time for one at the moment.

_edit_ Thinking a bit more about this, maybe a different exp is precisely the reason gcc -O3 was slower. I’ll try to benchmark just some pure functions like that later.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [November 21, 2018, 1:22pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/10 "2018-11-21T13:22:57Z")

</div>

> [@xor0110](#):
>
> I use functions like exp a lot in my code, so I’m definitely interesting in knowing how fast is my code with that, regardless of the reasons.

Including fast math? Enabling fast math is not a sign that one compiler or language is better than another. Among other things, it breaks IEEE compliance and only supports finite math. Yes, it can vastly speed up code, but I rarely find that I can actually use it in real world applications. Benchmarking one language with fast math enabled and one with it disabled (or semi-enabled), and drawing conclusions about general language performance, is nonsense IMO. If I take the Julia code in your article and replace the word `@simd` with `@fastmath`, the 1e9 test case goes from 17 seconds to 12 seconds on my system, which would completely obliterate any other benchmark in your article (assuming that you’re seeing the same timings, I haven’t tried your other implementations).

A few other notes:

- Julia can also be started with the `-O` flag to control optimization level. It defaults to 2; for optimum performance, consider setting this to 3 (doesn’t affect your code on my system, but it’s a good habit).
- Your Julia implementation seems to include a call to `abs` which is missing in the other languages?
- Prefer `System.nanoTime()` over `System.currentTimeMillis()` for benchmarking in JVM-based languages.
- The Python implementation using `%timeit` is unfair since it runs the code many times and selects the best time, while other languages just run it once. (I think it also disables GC, although that shouldn’t be an issue in your case since you’re not allocating memory.)

Hope I’m not coming across as too critical 🙂 I like your article, but accurate benchmarking is very difficult, so be careful making assumptions about results you get. Unexplainable results/anomalies are often caused by something wrong in the experiment itself, and not something on language-level.

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 21, 2018, 2:09pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/11 "2018-11-21T14:09:26Z")

</div>

I appreciate the scrutiny, I hope you do realize I just cooked up all of this code during the weekend, just looking for a first clue about whether it is true we can get high performance with Julia. There are certainly many ways it can be improved, this is by far not a mature benchmark! I am pretty aware of how difficult it is, but we need to start somewhere. I am just sharing my results as soon as possible to get more feedback earlier on. Please think of this more like a collaborative effort looking for contributors than a final external report that must be contradicted.

The experiments are completely clear about where fastmath was used, and I definitely hope I am in complete control of this all of the time, and when I say I don’t care why it is faster I do not mean using things like fastmath or even SIMD parallelism. I mean a possible faster implementation with a different method for exp that is equally accurate. As a good scientist should expect.

I have actually already heard from the C team that maybe Julia, Scala and Numpy could be all using fastmath implicitly for exp(), and that is the reason they might have been faster than “C -O3”. A pretty serious accusation, yet it is hard to prove. But we’ll get to the bottom of this eventually.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 21, 2018, 3:00pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/12 "2018-11-21T15:00:26Z")

</div>

> [@xor0110](#):
>
> I have actually already heard from the C team that maybe Julia, Scala and Numpy could be all using fastmath implicitly for exp(), and that is the reason they might have been faster than “C -O3”. A pretty serious accusation, yet it is hard to prove. But we’ll get to the bottom of this eventually.

Julia does not implicitly use fastmath. In fact, it uses its own Julia-based implementation in order to achieve ~1ulp since the system libraries do not always do so. So it has actually been shown that the Julia exp is more consistently accurate than the C stdlib versions, unless you add the `@fastmath` macro to change `exp` to `Base.FastMath.exp_fast`.

> [@xor0110](#):
>
> But we’ll get to the bottom of this eventually.

As native Julia code in an open source project, there is no “get to the bottom of this eventually”: you can get to the bottom of this right now by looking at the code yourself:

> <https://github.com/JuliaLang/julia/blob/master/base/special/exp.jl>

which is just standard Julia code. By clicking on the history for this file ( [History for base/special/exp.jl - JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/commits/master/base/special/exp.jl) ) you can see every edit and discussion that has currently gone into the development. As you can see from the PRs, this functionality is from @musm and comes from Amal.jl, his libm testing ground. In that repository you can actually see and run the benchmark which was used to verify the accuracy to 1ulp:

> <https://github.com/musm/Amal.jl/blob/master/benchmark/benchmark.jl>

This same setup was applied to research other Libms, like SLEEF (rewritten as Sleef.jl: [GitHub - musm/SLEEF.jl: A pure Julia port of the SLEEF math library](https://github.com/musm/SLEEF.jl) ), and was able to uncover inaccuracies in other libms such as [Log for small subnormals not accurate to within 1ulp · Issue #2 · musm/SLEEF.jl · GitHub](https://github.com/musm/SLEEF.jl/issues/2) .

So, knowing that all of this is online, what evidence did the C team give to state that Julia is implicitly using something equivalent to their ~3ulp fastmath?

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [November 21, 2018, 3:12pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/13 "2018-11-21T15:12:05Z")

</div>

Thanks a lot for your great answer. Indeed, looking into the Julia code has been the easiest and most enjoyable part of the investigation so far.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [November 21, 2018, 3:18pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/14 "2018-11-21T15:18:43Z")

</div>

Reading to code of Julia is almost always a pretty good idea, since it’s almost 70% pure Julia (as of today):

 ![image](https://global.discourse-cdn.com/julialang/original/3X/0/d/0dbb18aea2ae91be447f422bfcb9e95c2521b2e4.png)

So for learning how to do stuff in Julia it’s a good reference if you’re comfortable reading other people’s code.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 21, 2018, 4:37pm UTC](https://discourse.julialang.org/t/small-benchmark/17665/15 "2018-11-21T16:37:41Z")

</div>

> [@Sukera](#):
>
> since it’s almost 70% pure Julia (as of today):

Even that is a very low estimate too. The C and C++ is the Julia runtime and the Scheme is the Julia parser. In terms of Julia itself, what’s not implemented in Julia code is essentially the type system and its types like `DataType`, the `Array` type, and the expression types. The rest is all defined in Julia.

To show this, look at the top of boot.jl. This is the first file run on a Julia startup, and it comments on the top everything that exists in Julia prior to the Julia-defined Base code. This is 141 lines:

> <https://github.com/JuliaLang/julia/blob/master/base/boot.jl#L1-L141>

(though it leaves out a few primitive functions like `eval`). The rest is all defined in Julia, starting with integers and numbers:

> <https://github.com/JuliaLang/julia/blob/master/base/boot.jl#L179-L203>

The only caveat is the later portions of the stdlib defined by bindings to things like BLAS and SuiteSparse, or any packages you use which bind to binaries.

So seeing that, I think it’s fair to say that the Julia one actually interacts with is almost entirely written and defined in Julia, likely \>95%. The only pieces that people actually regularly touch that are define pre-Julia are the `Array`, `Union`, and `Expr` types. Even then, all of the functions on them, including the normal constructors, are defined in Julia.
