# Meaning of 2 minimum times of \`@btime\`

**URL:** <https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041>\
**Category:** General Usage\
**Tags:** benchmarktools\
**Created:** [June 20, 2022, 3:46am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041 "2022-06-20T03:46:18Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 3:46am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/1 "2022-06-20T03:46:18Z")

</div>

I know that \<1ns `@btime` timings mean that the compiler hoisted the method call out of the benchmark loop completely, but it’s a consistent number ~0.034ns. What does that number mean, where does it come from?

```julia
julia> @btime 1+1
  0.034 ns (0 allocations: 0 bytes)
2

julia> @btime 1+1+1+1
  0.034 ns (0 allocations: 0 bytes)
4

```

There’s another consistent minimum timing ~1.508ns, this time if I interpolate the values so the benchmark is doing something. I can’t make any benchmark go lower than this.

```julia
julia> @btime $1+$1
  1.508 ns (0 allocations: 0 bytes)
2

julia> @btime $1+$1+$1+$1
  1.508 ns (0 allocations: 0 bytes)
4

```

At first I thought one of these might be a single cpu cycle, but I’m doing this on a 3.1 GHz processor, and if my Googling is right, a cpu cycle is the reciprocal of that, 0.323ns, nowhere close to either figure.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [June 20, 2022, 4:14am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/2 "2022-06-20T04:14:01Z")

</div>

> [@Benny](#):
>
> a cpu cycle is the reciprocal of that, 0.323ns, nowhere close to either figure.

```julia
julia> @code_native debuginfo=:none (1+1)
	.text
	.file	"+"
	.globl	"julia_+_847" # -- Begin function julia_+_847
	.p2align	4, 0x90
	.type	"julia_+_847",@function
"julia_+_847": # @"julia_+_847"
	.cfi_startproc
# %bb.0: # %top
	leaq	(%rdi,%rsi), %rax
	retq
.Lfunc_end0:
	.size	"julia_+_847", .Lfunc_end0-"julia_+_847"
	.cfi_endproc
                                        # -- End function
	.section	".note.GNU-stack","",@progbits

```

this doesn’t look like 1 cycle to me.

In general, O(ns) benchmarking is not meaningful / reliable, and also modern CPU has pipelines so “a single cycle / instruction” is not as clean-cut as you might think

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 7:04am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/3 "2022-06-20T07:04:55Z")

</div>

But I would normally expect something unreliable to vary randomly. These are consistent 0.034ns and 1.508ns I’m seeing. I can accept that these are beyond the limits of timing precision and cannot mean anything close to “this is how long this method call took”, but they must come from somewhere.

And why doesn’t `@btime` only report meaningful timings i.e. never go below a timing known to be the limit of meaningfulness? Is 0.001ns even a meaningful precision for any subrange of 0-10ns timings?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [June 20, 2022, 8:40am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/4 "2022-06-20T08:40:53Z")

</div>

My gut says these are probably related to the instruction level caches your CPU does. There are more caches than just L1, L2, etc., so I’m guessing “return this constant” is caching the constant itself, as well as where to put it - with potential microcode caching & optimizations as well. Those are no longer transparent to BenchmarkTools.

> [@Benny](#):
>
> And why doesn’t `@btime` only report meaningful timings i.e. never go below a timing known to be the limit of meaningfulness? Is 0.001ns even a meaningful precision for any subrange of 0-10ns timings?

It’s just what the high precision timer spits out. IIRC there was a PR/issue for it about warning when such small numbers are returned, precisely to make people aware that those sub nanosecond timings are bogus.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [June 20, 2022, 10:29am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/5 "2022-06-20T10:29:20Z")

</div>

I don’t think this was mentioned already, but subnanosecond timings should become a memory in Julia v1.8, thanks to [this PR](https://github.com/JuliaCI/BenchmarkTools.jl/pull/275) which makes use of the new `Base.donotdelete`:

```julia
julia> @btime 1 + 1
  1.370 ns (0 allocations: 0 bytes)
2

```

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 10:36am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/6 "2022-06-20T10:36:42Z")

</div>

So what does `Base.donotdelete` do? Seems odd this is a part of Base rather than something entirely in BenchmarkTools. Is that some way to tell the compiler “don’t hoist”?

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [June 20, 2022, 10:44am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/7 "2022-06-20T10:44:46Z")

</div>

It prevents [dead-code elimination](https://en.wikipedia.org/wiki/Dead-code_elimination). This makes the benchmarks somewhat less realistic in the case the compiler _is_ able to constant-propagate the result, but it also makes them more useful, as you get a better lower bound of the performance you can achieve without hoping the compiler will optimise everything away.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 10:50am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/8 "2022-06-20T10:50:05Z")

</div>

Digging into the compiler seems drastic. The way it was explained to me, hoisting meant it was effectively benchmarking `() -> 1+1`. But shouldn’t it be doable to figure out from `@btime 1+1` that `(x, y) -> x+y` is the appropriate benchmark, and `x, y = 1, 1` just happens to be true every iteration? Or is this not the correct explanation?

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [June 20, 2022, 10:56am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/9 "2022-06-20T10:56:03Z")

</div>

> [@Benny](#):
>
> But shouldn’t it be doable to figure out from `@btime 1+1` that `(x, y) -> x+y` is the appropriate benchmark

Is it?

```julia
julia> @btime ((x, y) -> x + y)(1, 1)
  0.012 ns (0 allocations: 0 bytes)
2

```

(julia v1.7)

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 11:07am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/10 "2022-06-20T11:07:12Z")

</div>

The way it was explained to me, `@btime ((x, y) -> x + y)(1, 1)` would be benchmarking `() -> ((x, y) -> x + y)(1, 1)`, which is compiled to `() -> 2`. Basically the benchmarked expression is put in a function body, and how many inputs you interpolate is how many arguments that wrapper function gets.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [June 20, 2022, 11:16am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/11 "2022-06-20T11:16:10Z")

</div>

Right. The reason you need Base.donotdelete is to percent the compiler from constant proping away the wrapper.

---

<div class="post-metadata">

**Author:** ![albheim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/albheim/32/34660_2.png) [@albheim](https://discourse.julialang.org/u/albheim)\
**Post date:** [June 20, 2022, 11:37am UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/12 "2022-06-20T11:37:27Z")

</div>

I think what @Benny is saying is that maybe `BenchmarkTools` could automatically be able to “extract” constants from the expression, and instead of benchmarking `1+1`, which is compiled away, it benchmarks `x+y` for some values. That the actual benchmark happens with the values `1,1` should then not have an effect on the compiled function, but is something that happens when the benchmark calls the function.

```julia
julia> f(x, y) = x + y
f (generic function with 1 method)

julia> @code_llvm f(1,1)
; @ REPL[3]:1 within `f`
define i64 @julia_f_1012(i64 signext %0, i64 signext %1) #0 {
top:
; ┌ @ int.jl:87 within `+`
   %2 = add i64 %1, %0
; └
  ret i64 %2
}

julia> @btime f(1, 1) # This is removed by compiler
  0.026 ns (0 allocations: 0 bytes)
2

julia> @btime f(x, y) setup=begin x=1; y=1 end # This I expected to work
  0.026 ns (0 allocations: 0 bytes)
2

julia> @btime f(x, y) setup=begin x=rand([1]); y=rand([1]) end # This works
  1.952 ns (0 allocations: 0 bytes)
2

```

So the first one is as expected compiled away. The second I expected to work, but it also seems to be compiled away. The third seems to do the trick though, seems like it runs some instructions at least.

So if this is now possible, could `BenchmarkTools` automatically do this extraction of constants, and set them up in a similar way as above?

This might be the same (or at least similar) to what is done in 1.8?

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 20, 2022, 5:42pm UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/13 "2022-06-20T17:42:37Z")

</div>

Okay I was under the very wrong impression that `@btime` is making the “wrapper function” from the input expression; it’s the compiler doing its usual optimizations (constant propagation, loop-invariant code motion) to the benchmark loop, hence the utility of having a way to tell the compiler to not do it.

A less confused question: we _can_ manually `$`-interpolate to get the behavior we want, so why resort to the new `Base.donotdelete` instead of making `@btime` automatically insert `$` before arguments in the input expression? It just seems like the latter approach can also benefit past versions of Julia. I suppose it won’t cover the Ref trick, but that seems to be less necessary now (at least not anymore for the example in the docs’ Quick Start).

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [June 20, 2022, 7:25pm UTC](https://discourse.julialang.org/t/meaning-of-2-minimum-times-of-btime/83041/14 "2022-06-20T19:25:08Z")

</div>

To get back to the original question on fixed times, the minimum delta that `time_ns()` can measure seems to be limited (and is likely system-dependent). On my system:

```julia
julia> function f()
           x, y = time_ns(), time_ns()
           return Int(y - x)
       end
f (generic function with 1 method)

julia> first(sort(unique(f() for i in 1:1000000)), 10)
10-element Vector{Int64}:
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41

julia> @benchmark 1+1
BenchmarkTools.Trial: 10000 samples with 1000 evaluations.
 Range (min … max): 0.034 ns … 0.088 ns ┊ GC (min … max): 0.00% … 0.00%

```

Note that it’s grouping the execution into groups of 1000 evaluations. `34/1000 == 0.034`. It’s not as obvious to me where your regularly appearing 1.508ns stems from, but IIRC when you interpolate Julia does a function call.
