# Making @benchmark outputs statistically meaningful, and actionable

**URL:** <https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256>\
**Category:** Profiling\
**Tags:** benchmark, benchmarktools\
**Created:** [September 26, 2023, 11:38am UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256 "2023-09-26T11:38:25Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 11:38am UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/1 "2023-09-26T11:38:25Z")

</div>

I see a huge disparity in the output of different runs of `@benchmark`, for the same code. See the following two outputs, made consecutively, with zero code changes and no restarting Julia:

```julia
BenchmarkTools.Trial: 161 samples with 1 evaluation.
 Range (min … max): 30.098 ms … 32.380 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 30.904 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 31.217 ms ± 632.639 μs ┊ GC (mean ± σ): 0.00% ± 0.00%

                ▅█▅█▂█▆ ▂ ▃▅ ▂▃▅      
  ▄▄▁▄▅▄▅▄▇▄▅▄▇▅███████▅█▁▇▄▁▄▄▁▄▁▄▁▁▄▁▁▁▅▅▅▇▅██▇▇▄▅▇█▇███▅▁▅▇ ▄
  30.1 ms Histogram: frequency by time 32.2 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

BenchmarkTools.Trial: 119 samples with 1 evaluation.
 Range (min … max): 40.943 ms … 44.543 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 41.448 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 42.238 ms ± 1.218 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

    ▃ ▂ ▅█▂                                                    
  ▆▅█▇█▆███▄▄▁▃▃▅▃▃▁▃▁▁▁▁▃▄▁▁▃▁▁▁▁▁▁▁▃▁▁▄▄▄▄▄▃▁▃▅▁▅▅▃▃▃▅▄▇▅▃▄ ▃
  40.9 ms Histogram: frequency by time 44.4 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

Note there is no garbage collection at play here.

There is a _huge_ difference in the two outputs. There isn’t any overlap between the two histograms. In fact, there’s around a 30% difference between the slowest time in the first run and the fastest time in the next run! It’s clear that the histograms are therefore meaningless – they don’t reflect the real distribution of timings.

This makes results from `@benchmark` lack actionability. If I want to know whether a small code change leads to faster code, `@benchmark` cannot reliably help me answer that. I might randomly get a faster run for code that is in general slower.

So how is this happening?

I believe that this issue is due to alignment between working memory and cache lines. Different runs of the code will use memory blocks at different offsets from the start of a cache line. This can have a large impact on cache hits/misses and therefore performance. Usually we have no control over this at runtime (though it’s possible to ensure that memory blocks have a particular alignment if needed).

However, there’s a fix – we can make sure that the results from `@benchmark` reflect this variability. In order to do this, it needs to randomly change the offset of stack and heap allocated memory between evaluations of the code.

Indeed, we can even go further, and say whether the difference between two runs is statistically significant. For this we’d need a new macro, `@bcompare`.

About 3-4 years ago I watched a video of great talk by some CS Professor who explains this and has developed a benchmarking tool that accounts for it (not for Julia, for command line applications). ~~Wish I could find it again, but I haven’t been able to work out the right keywords to find it in Google. If you know what I’m talking about, please can you share a link?~~ Well worth a watch, [here](https://www.youtube.com/watch?v=r-TLSBdHe1A).

---

<div class="post-metadata">

**Author:** ![ArchieCall](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/archiecall/32/206_2.png) [@ArchieCall](https://discourse.julialang.org/u/ArchieCall)\
**Post date:** [September 26, 2023, 11:46am UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/2 "2023-09-26T11:46:12Z")

</div>

What is MWE. Need that to give any meaningful answer.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 11:49am UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/3 "2023-09-26T11:49:36Z")

</div>

Answer to what, Archie? The post consists of statements, not questions. Well, there is one question, about a link to a talk, but I don’t think that requires a MWE.

The code I’m benchmarking is reasonably complex, and reducing it down to a MWE would be a lot of effort. At this point it’s not clear what use that MWE would serve. Do you not believe me about the timings? Have you never seen something similar?

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [September 26, 2023, 12:00pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/4 "2023-09-26T12:00:03Z")

</div>

Some type of of MWE would indeed be helpful here beucase although there are no code questions being asked, recommendations are being made to change/update/create new benchmarking tools. But before people agree that there is a problem, they need to be able to reproduce it themselves. I myself have seen somehwat varying timings, but usually I can trace the results to a “me problem” rather than a problem with `BenchmarkTools`.

Some questions I would like to test are, what if `@benchmark` is 5 times in a row? What about after a fresh session restart? Does it change if you use randomized inputs? What about setting the `seconds` keyword to increase the number of runs? Do the results come out the same? If I can’t test these questions with a MWE, then I can’t tell if there is a problem unless I am somehow able to write my own code with unstable benchmarks.

And its not about believing or not believing, its about figuring out where the issue lies before recommending code changes.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 12:02pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/5 "2023-09-26T12:02:47Z")

</div>

So I found [the talk](https://www.youtube.com/watch?v=r-TLSBdHe1A). Well worth a watch.

---

<div class="post-metadata">

**Author:** ![devel-chm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/devel-chm/32/3572_2.png) [@devel-chm](https://discourse.julialang.org/u/devel-chm)\
**Post date:** [September 26, 2023, 1:10pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/6 "2023-09-26T13:10:39Z")

</div>

Two thoughts:

1. I’ve found that benchmarks for short duration functions  
can be tricky and mysteriously variable due to all the hardware  
and OS and runtime thing that can affect operations.

2. I notice that the total times of both benchmark runs are about  
4991 msec for the first and 4939 msec but the number of  
samples are 161 and 119 respectively. All of the difference  
between the two appears to be from the sample counts.

Maybe understanding why the number of samples are different  
between runs could explain the timings or suggest a MWE?

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 1:13pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/7 "2023-09-26T13:13:13Z")

</div>

> [@devel-chm](#):
>
> Maybe understanding why the number of samples are different  
> between runs could explain the timings or suggest a MWE?

Unfortunately not. @benchmark is simply trying to run evaluations for 5 seconds total. The number of samples is inversely proportional to the function evaluation time.

---

<div class="post-metadata">

**Author:** ![devel-chm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/devel-chm/32/3572_2.png) [@devel-chm](https://discourse.julialang.org/u/devel-chm)\
**Post date:** [September 26, 2023, 1:27pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/8 "2023-09-26T13:27:41Z")

</div>

Ok.

Have you tried running the benchmark repeatedly for an extended  
time to collect overall variability?

Is it possible that the processor clock changed between the runs?

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 1:39pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/9 "2023-09-26T13:39:16Z")

</div>

> [@devel-chm](#):
>
> Have you tried running the benchmark repeatedly for an extended  
> time to collect overall variability?

The _whole_ point of `@benchmark` is that a single run shows the overall variability. That it doesn’t is precisely the problem I’m trying to highlight.

> [@devel-chm](#):
>
> Is it possible that the processor clock changed between the runs?

It isn’t.

Have you watched the video? Memory layout affects runtime. `@benchmark` doesn’t account for this. I’m not saying anything controversial.

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [September 26, 2023, 1:50pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/10 "2023-09-26T13:50:42Z")

</div>

This seems like an interesting proposition, despite a slightly confrontational tone 😉  
Do you know how to implement what you’re suggesting? I’m basically the sole maintainer of BenchmarkTools.jl but I’m not sure I’m knowledgeable enough about low-level memory stuff.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 2:14pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/11 "2023-09-26T14:14:11Z")

</div>

> [@gdalle](#):
>
> despite a slightly confrontational tone 😉

A habit of mine. Especially when people suggest or imply that I’ve simply made a mistake. Not that it’s impossible, mind.

> [@gdalle](#):
>
> Do you know how to implement what you’re suggesting?

My thought is to allocate a randomly sized (within the length of a cache line, usually 64/128 bytes) block on the stack, and another larger, but randomly sized, block on the heap between evaluations. This would be a simple hack.

However, I note that in the talk, Emery says (at 22:00) their tool ‘Stabilize’ directly plugs in to the llvm compiler, so it could potentially be used directly. This would be the ideal, since it does far more randomization. For example, if someone uses a test function that generates the data for benchmarking, the heap allocation layout will usually be the same, even with a slight offset. You really want something that jumbles everything up.

I could take a look at the BenchmarkTools code and try to do the former; see if it increases the variability as expected. I’ll also see if I can create a MWE. The problem is that the effect is hardware and OS dependent, so it might not reproduce in everyone’s environment.

---

<div class="post-metadata">

**Author:** ![devel-chm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/devel-chm/32/3572_2.png) [@devel-chm](https://discourse.julialang.org/u/devel-chm)\
**Post date:** [September 26, 2023, 2:14pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/12 "2023-09-26T14:14:28Z")

</div>

> The _whole_ point of `@benchmark` is that a single run shows the overall variability. That it doesn’t is precisely the problem I’m trying to highlight.

I was wondering what the distribution of benchmark results was  
compared to the per-run results and statistics to understand exactly  
how often and by how much this difference occurs.

> [@devel-chm](#):
>
> Is it possible that the processor clock changed between the runs?

> It isn’t.
> 
> Have you watched the video? Memory layout affects runtime. `@benchmark` doesn’t account for this. I’m not saying anything controversial.

I’m not sure `@benchmark` corrects for varying CPU clock speed so  
I wanted to confirm that this was not the issue.

I understand how memory layout can affect runtime and don’t think you  
have been saying anything controversial.

_I don’t know how to get the quoted text with user name in the link.  
Apologies for that._

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [September 26, 2023, 3:13pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/13 "2023-09-26T15:13:16Z")

</div>

> [@user664303](#):
>
> Especially when people suggest or imply that I’ve simply made a mistake.

Rereading the posts, I don’t think anyone suggested or implied mistakes were made. I commented on my experience doing benchmarks, but I was simply answering your question asking “Have you never seen something similar?”. To which is the answer is “no”, I have not experienced variance that I couldn’t trace back to an issue of my own.

All the other posts are simply asking questions to obtain more knowledge about the problem since I/we do not have code that reproduces the issue, so we are relying on the information you provide to better understand what is happening. If the problem is statistical in nature (results change depending on circumstance), then it seems a sample size larger than one is necessary to properly understand it.

I’m also not saying-- or implying – that you’re wrong, I certainly have no experience in this area to make that judgement. But when someone makes a post showing incorrect/inconsistent results, the first step taken by those trying to help is almost universally: try to reproduce the results. That is all we are asking.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [September 26, 2023, 3:37pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/14 "2023-09-26T15:37:24Z")

</div>

Have you tried passing a `setup` block that re-allocates/copies/shuffles around your data?

It could be caching/locality effects, but it could also be _so many things_. Modern computer architectures (and operating systems) are wild. From thermal throttling to [(surprisingly long-lived)](https://discourse.julialang.org/t/psa-microbenchmarks-remember-branch-history/17436) branch prediction to caching to multithreading to hyperthreading… you’re performance fine-tuning in a hostile environment.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [September 26, 2023, 3:43pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/15 "2023-09-26T15:43:09Z")

</div>

Providing a MWE is the first step in collectively solving a problem here. See the post pinned to the top of the forum when you first land here:

> [@Please read: make it easier to help you](https://discourse.julialang.org/t/please-read-make-it-easier-to-help-you/14757):
>
> Welcome to the Julia Discourse! We are enthusiastic about helping Julia programmers, both beginner and experienced. This public service announcement (PSA) outlines best practices when asking for help. Following these points makes it easier for us to help you and more likely you’ll get a prompt, useful answer. Keywords are highlighted to make it easier to refer to specific points. Choose a descriptive title that captures the key part of your question, eg “plots with multiple axes” instead of …

Someone asking for a MWE should not at all be taken as offensive to you at all - it’s the expected request on this forum if you haven’t included one in your first post.

Without an MWE you are asking everyone to reproduce your work by guessing, instead of diving in and trying to understand your problem straight away.

You will get way more help and attention on any topic if you have a MWE.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [September 26, 2023, 4:07pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/16 "2023-09-26T16:07:23Z")

</div>

```julia
julia> f() = sleep(rand(1:10))
f (generic function with 1 method)

julia> using BenchmarkTools

julia> @benchmark f()
BenchmarkTools.Trial: 1 sample with 1 evaluation.
 Single result which took 10.008 s (0.00% GC) to evaluate,
 with a memory estimate of 112 bytes, over 4 allocations.

julia> @benchmark f()
BenchmarkTools.Trial: 2 samples with 1 evaluation.
 Range (min … max): 2.001 s … 10.008 s ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 6.004 s ┊ GC (median): 0.00%
 Time (mean ± σ): 6.004 s ± 5.662 s ┊ GC (mean ± σ): 0.00% ± 0.00%

  █ █  
  █▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ ▁
  2 s Histogram: frequency by time 10 s <

 Memory estimate: 112 bytes, allocs estimate: 4.

```

Unless we know anything about the code being benchmarked the results on their own are not surprising.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 6:30pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/17 "2023-09-26T18:30:05Z")

</div>

I don’t think you did say I was wrong. Or that you implied it. Nothing I said implied that. I can see why you inferred it. It’s easy to infer things that aren’t true.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 6:32pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/18 "2023-09-26T18:32:28Z")

</div>

Again, I never said anybody said anything offensive here.

I did say that an MWE would be tricky, for two different reasons. Did you notice?

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 6:37pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/19 "2023-09-26T18:37:12Z")

</div>

My results are different from this. And you do know something about the code: 1. The code was run twice in succession. 2. It doesn’t allocate any memory.

I can tell you more:  
3. It’s pure computation. No disk IO, no network IO, no screen output.

I would think that would make the results surprising.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [September 26, 2023, 6:44pm UTC](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/20 "2023-09-26T18:44:40Z")

</div>

Yes! And we account for most of them in the benchmark 😄 If we want to.

I will look into your setup block suggestion.

[Next page](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256.md?page=2)
