# How to benchmark properly? Should defaults change?

**URL:** <https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138>\
**Category:** General Usage\
**Tags:** benchmarktools\
**Created:** [August 10, 2021, 2:44pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138 "2021-08-10T14:44:59Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 10, 2021, 2:44pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/1 "2021-08-10T14:44:59Z")

</div>

What is correct here? I guess is the one with `evals=1`. In that case, couldn’t it be the default option?

```julia
julia> using BenchmarkTools

julia> @btime sin(5.0)
  1.538 ns (0 allocations: 0 bytes)
-0.9589242746631385

julia> x = 5.0
5.0

julia> @btime sin($x)
  7.530 ns (0 allocations: 0 bytes)
-0.9589242746631385

julia> @btime sin($x) evals=1
  40.000 ns (0 allocations: 0 bytes)
-0.9589242746631385

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 10, 2021, 2:47pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/2 "2021-08-10T14:47:13Z")

</div>

`@btime` gives you the minimal time, so of course `eval=1` will make a difference

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [August 10, 2021, 2:50pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/3 "2021-08-10T14:50:03Z")

</div>

When things can be constant propagated, I usually `Ref` wrap the input like:

```julia
julia> using BenchmarkTools

julia> @btime sin(5.0) # bogus result from const prop
  0.013 ns (0 allocations: 0 bytes)
-0.9589242746631385

julia> x = 5.0
5.0

julia> @btime sin($x) # bogus result from const prop
  0.013 ns (0 allocations: 0 bytes)
-0.9589242746631385

julia> @btime sin($(Ref(x))[]) # ok
  6.476 ns (0 allocations: 0 bytes)
-0.9589242746631385

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 10, 2021, 2:51pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/4 "2021-08-10T14:51:59Z")

</div>

I don’t think is simply that. With `evals=1` it is running many samples of 1 evaluation each. Without it, it is running many samples of many evaluations (at least is what I understand from the manual). My understanding is that when more than one evaluation per sample is being run, we are getting some artifact associated to caching results.

```julia
julia> @benchmark sin($x)
BenchmarkTools.Trial: 10000 samples with 999 evaluations.
 Range (min … max): 7.539 ns … 42.284 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 8.425 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 9.240 ns ± 1.905 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

      █ ▅                                                     
  ▁▂▆▅███▁▁▁▂▂▃▁▃▁▂▁▂▁▂▅▁▆▂▁▄▁▁▂▁▁▁▂▁▁▁▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ ▂
  7.54 ns Histogram: frequency by time 15.4 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sin($x) evals=1
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 41.000 ns … 542.000 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 45.000 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 45.989 ns ± 5.730 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

           ▁ ▆ █                                          
  ▃▁▁▁▅▁▁▁▁█▁▁▁▁█▁▁▁▁█▁▁▁▁▇▁▁▁▁▆▁▁▁▁▇▁▁▁▁▅▁▁▁▁▅▁▁▁▁▄▁▁▁▁▄▁▁▁▁▃ ▃
  41 ns Histogram: frequency by time 53 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

@kristoffer.carlsson I think you still have constant propagation in your `Ref` example, but within samples. Might be wrong though. What you get if you use `evals=1`? (your timings are very different from mine here).

Here I get:

```julia
julia> @btime sin($(Ref(x))[]) 
  7.778 ns (0 allocations: 0 bytes)
-0.9589242746631385

julia> @btime sin($(Ref(x))[]) evals=1
  40.000 ns (0 allocations: 0 bytes)
-0.9589242746631385

```

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [August 10, 2021, 2:55pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/5 "2021-08-10T14:55:27Z")

</div>

> [@lmiq](#):
>
> (your timings are very different from mine here).

I use nightly Julia which probably has better constant propagation.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 10, 2021, 3:00pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/6 "2021-08-10T15:00:08Z")

</div>

That is good news for Julia, but not so for `BenchmarkTools`, which becomes harder to understand.

I still don’t understand the difference in the results. It is even strange that, for example, using `evals=1` one gets:

```julia
julia> @btime sin($x) evals=1
  40.000 ns (0 allocations: 0 bytes)
-0.9589242746631385

```

and with `evals=2` one gets:

```julia
julia> @btime sin($x) evals=2
  24.500 ns (0 allocations: 0 bytes)
-0.9589242746631385

```

And by increasing `evals` one converges to about `7 ns` which is (?) the correct benchmark (?).

These are completely systematic, thus these results do not seem to be associated with random fluctuations of the benchmark.

---

<div class="post-metadata">

**Author:** ![tisztamo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tisztamo/32/16200_2.png) [@tisztamo](https://discourse.julialang.org/u/tisztamo)\
**Post date:** [August 10, 2021, 3:02pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/7 "2021-08-10T15:02:04Z")

</div>

I checked the code a few months ago and it seemed that the runtime of a single `time_ns()` call is always added to the measuerement, resulting in a 25ns/evals error on my machine.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 10, 2021, 3:04pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/8 "2021-08-10T15:04:58Z")

</div>

Ah, that is one reason for the problem. Ok. So that systematic error is diluted when one uses many evaluations in each sample.

That mixed with the constant propagation thing makes makes benchmarking a little bit confusing. Perhaps there is room for improvement in the API?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [August 10, 2021, 3:12pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/9 "2021-08-10T15:12:10Z")

</div>

I’ve found that I get more robust results by just broadcasting over an input vector:

```julia
julia> @btime sin.(x) setup=(x=rand(1000));
  6.083 μs (1 allocation: 7.94 KiB)

```

It’s less ergonomic, but it does a good job guarding against overly-aggressive constant propagation.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [August 10, 2021, 8:23pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/10 "2021-08-10T20:23:14Z")

</div>

This is also good for functions with branches. For example, sin will be faster for numbers less than pi/4.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [August 10, 2021, 9:06pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/11 "2021-08-10T21:06:05Z")

</div>

Plus, it makes the SIMD implementations look good. 😉

I normally `Ref`-wrap any `isbits` structs I’m benchmarking.

40ns is way too long. That’d be well over 100 clock cycles for most CPUs.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 10, 2021, 9:09pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/12 "2021-08-10T21:09:05Z")

</div>

> [@Elrod](#):
>
> I normally `Ref` -wrap any `isbits` structs I’m benchmarking

Couldn’t that be automatic?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [August 10, 2021, 9:11pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/13 "2021-08-10T21:11:54Z")

</div>

It’s of course possible that the compiler will get smart enough to defeat this, too, eventually.

But I’d be in favor of it ref-wrapping everything by default.

The macro doesn’t have access to type information, but it could probably generate code that’s the equivalent of

```julia
rx = Ref(x)
isbits(x) ? x : rx[]

```

Or maybe just ref-wrap everything by default.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [August 10, 2021, 9:48pm UTC](https://discourse.julialang.org/t/how-to-benchmark-properly-should-defaults-change/66138/14 "2021-08-10T21:48:27Z")

</div>

There is some disagreement on whether minimum should be used.

[Robust benchmarking in noisy environments](https://arxiv.org/abs/1608.04295) (2016)

[Minimum Times Tend to Mislead When Benchmarking](https://tratt.net/laurie/blog/entries/minimum_times_tend_to_mislead_when_benchmarking.html) (2019)
