# Bad performance from parametric struct dispatch

**URL:** <https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594>\
**Category:** Performance\
**Created:** [July 3, 2024, 3:00pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594 "2024-07-03T15:00:27Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![quantumtwist](https://avatars.discourse-cdn.com/v4/letter/q/ebca7d/32.png) [@quantumtwist](https://discourse.julialang.org/u/quantumtwist)\
**Post date:** [July 3, 2024, 3:00pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/1 "2024-07-03T15:00:27Z")

</div>

While trying to optimize the hot loop of my code, I ran into a very strange performance footgun that is reducing the speed of a simple function by at least ~100ns, which is a 10x performance hit for my use case.

**Question:** Why do the two almost-identical functions below have such different performance?

**Background:** I am working with binary configurations over a small grid (e.g. 3x7 or 6x8). Configurations are encoded in the low bits of a `UInt64` or similar. Since the grid size is known in advance, I use a parametric struct `grid = Grid{N1,N2}` to represent it, and pass the grid to functions so they can specialize on it (since this is in the hot loop). For example: `do_something(data, grid :: Type{Grid{N1,N2}) where {N1,N2}`. However, this seems to take ~100ns longer than `do_something(data, ::Val{N1}, ::Val{N2}) where {N1,N2}`. I would have thought that, since `N1` and `N2` are known at compile time, these approaches would be completely equivalent, but the second is almost 10x faster. What’s going on here? Is there a way to make something in the style of the first function fast? (Also, are there any other ways to speed my function up?)

**Minimal Working Example:**

```julia
struct Grid{N1,N2}
    #other data here
end

struct GridInd{G <: Grid}
    n1 :: Int8
    n2 :: Int8
    function GridInd{Grid{N1,N2}}(n1, n2) where {N1,N2}
        return new{Grid{N1,N2}}(mod(n1,N1),mod(n2,N2))
    end
end
#just for display
showconfig(config, gd :: Type{Grid{N1,N2}}) where {N1,N2} = reverse(reshape(digits(config,base=2,pad=N1*N2),N1,N2),dims=2)'

function momentum_v1(config, ::Type{Grid{N1,N2}}) where {N1,N2}
    K1, K2 = 0, 0
    while config != 0
        place = trailing_zeros(config) #place of last set 1
        k2, k1 = fldmod(place,N1)
        K1 += k1
        K2 += k2
        config = config & (config-1) #remove last set 1
    end
    return GridInd{Grid{N1,N2}}(K1,K2)
end

#only the function signature is different than v1
function momentum_v2(config, ::Val{N1}, ::Val{N2}) where {N1,N2}
    K1, K2 = 0, 0
    while config != 0
        place = trailing_zeros(config) #place of last set 1
        k2, k1 = fldmod(place,N1)
        K1 += k1
        K2 += k2
        config = config & (config-1) #remove last set 1
    end
    return GridInd{Grid{N1,N2}}(K1,K2)
end

grid = Grid{3,7}
config = UInt64(4325) #typical example configuration
display(showconfig(config,grid)) #show configuration for clarity

using BenchmarkTools
@benchmark momentum_v1($config,$grid)
v1,v2= Val(3),Val(7)
@benchmark momentum2($config,$v1,$v2) #fixed interpolations

```

This yields:

```julia
#example configuration
7×3 adjoint(::Matrix{Int64}) with eltype Int64:
 0 0 0
 0 0 0
 1 0 0
 0 0 0
 1 1 0
 0 0 1
 1 0 1

# momentum_v1
BenchmarkTools.Trial: 10000 samples with 925 evaluations.
 Range (min … max): 112.027 ns … 514.099 ns ┊ GC (min … max): 0.00% … 72.97%
 Time (median): 113.018 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 115.742 ns ± 11.667 ns ┊ GC (mean ± σ): 0.18% ± 1.75%

  ▅█▅▂▃▃▄▃▁▂▁▁ ▁
  █████████████▇██▇▆▆▆▇▆▇▆▇▇▆▆▇▇▇▇▆▇▇▇▇▆▇▆▆▆▆▆▆▅▅▅▆▅▅▅▆▅▅▄▄▄▅▅▄ █
  112 ns Histogram: log(frequency) by time 147 ns <

 Memory estimate: 32 bytes, allocs estimate: 2.

# momentum_v2
BenchmarkTools.Trial: 10000 samples with 1000 evaluations.
 Range (min … max): 10.333 ns … 54.792 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 10.417 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 10.588 ns ± 1.594 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

   █                                                           
  ▃█▄▂▂▂▂▂▁▁▂▂▂▂▁▁▁▂▂▁▁▁▁▂▂▁▁▁▁▁▁▁▁▁▁▁▂▁▁▁▁▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂ ▂
  10.3 ns Histogram: frequency by time 14.7 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

---

<div class="post-metadata">

**Author:** ![brianguenter](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brianguenter/32/29519_2.png) [@brianguenter](https://discourse.julialang.org/u/brianguenter)\
**Post date:** [July 3, 2024, 3:24pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/2 "2024-07-03T15:24:21Z")

</div>

Don’t know why this is happening but I did notice that the second argument `v2` is not prefixed by $ in one of the benchmarks, which I assume is what you intended to do. I made this change and the benchmark ran a little faster.

> [@quantumtwist](#):
>
> ```julia
> v1,v2= Val(3),Val(7)
> @benchmark momentum2($config,$v1,v2)
> 
> ```

I looked at the code using @code\_warntype for both functions. They generate nearly identical code, nothing stands out as being obviously inefficient. Maybe a compiler guru can explain what’s going on.

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [July 3, 2024, 3:27pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/3 "2024-07-03T15:27:02Z")

</div>

I think you should avoid the argument `::Type{Grid{N1,N2}}`, it requires some allocations for dispatch. If you instead use `momentum_v1(config, ::Grid{N1,N2})`, there are no allocations, and it’s as fast as the `Val` version, though you must give it a `Grid` object, not the type.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [July 3, 2024, 3:32pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/4 "2024-07-03T15:32:49Z")

</div>

[https://docs.julialang.org/en/v1/manual/performance-tips/#Be-aware-of-when-Julia-avoids-specializing](https://docs.julialang.org/en/v1/manual/performance-tips/#Be-aware-of-when-Julia-avoids-specializing)

---

<div class="post-metadata">

**Author:** ![quantumtwist](https://avatars.discourse-cdn.com/v4/letter/q/ebca7d/32.png) [@quantumtwist](https://discourse.julialang.org/u/quantumtwist)\
**Post date:** [July 3, 2024, 3:37pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/5 "2024-07-03T15:37:59Z")

</div>

[brianguenter](https://discourse.julialang.org/u/brianguenter) Thank you for catching that typo. My benchmarks above used the proper interpolation for v2.

[lmiq](https://discourse.julialang.org/u/lmiq) The page you linked seems very helpful. It claims

```julia
#but this will [specialize]

function g_type(t::Type{T}) where T
    x = ones(T, 10)
    return sum(map(sin, x))
end

```

which I thought was the case I am in here.

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [July 3, 2024, 3:38pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/6 "2024-07-03T15:38:06Z")

</div>

After looking into this. It’s actually an artifact of the benchmarking, not an actual slowdown in a real program.

```julia
@benchmark momentum_v1($config,$(Ref{Type{Grid{3,7}}}(grid))[])

```

is fine.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [July 3, 2024, 3:49pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/7 "2024-07-03T15:49:31Z")

</div>

To give another way of avoiding the benchmarking artifacts:

```julia
julia> foo1(config) = momentum_v1(config, Grid{3,7});

julia> foo2(config) = momentum_v2(config, Val(3), Val(7))

julia> @btime foo1($config)
  7.450 ns (0 allocations: 0 bytes)
GridInd{Grid{3, 7}}(2, 2)

julia> @btime foo2($config)
  7.438 ns (0 allocations: 0 bytes)
GridInd{Grid{3, 7}}(2, 2)

```

The exact details of the benchmarking loop body is somewhat of a black magic. Generally, if results turn out surprising, I would advise to write your own benchmarking loop. That way, you control inlining and what is known to the compiler, and you know whether you’re benchmarking throughput or also latency. Big difference whether you need to apply a function to lots of elements independently (throughput!) or the next function application needs the result of the last as input (latency!). You also can guess at the state of caches, branchpredictor, etc and adjust to more closely match your expected real-world state.

---

<div class="post-metadata">

**Author:** ![quantumtwist](https://avatars.discourse-cdn.com/v4/letter/q/ebca7d/32.png) [@quantumtwist](https://discourse.julialang.org/u/quantumtwist)\
**Post date:** [July 3, 2024, 3:50pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/8 "2024-07-03T15:50:34Z")

</div>

While that speeds up the microbenchmark, the performance I get from putting `momentum` inside another function is quite different between the two cases.

```julia
using Random
function do_it_many_times(N,grid)
    a = 0
    Random.seed!(4432543)
    for n in 1:N
        config = rand(UInt64) & (1<<22 -1)
        K = momentum_v1(config,grid)
        a += K.n1 #prevent compiler from trivializing loop
    end
    return a
end

function do_it_many_times2(N,v1,v2)
    a = 0
    Random.seed!(4432543)
    for n in 1:N
        config = rand(UInt64) & (1<<22 -1)
        K = momentum_v2(config,v1,v2)
        a += K.n1 #prevent compiler from trivializing loop
    end
    return a
end

N = 1000
gd = Grid{3,7}
@btime do_it_many_times($N,$gd)
@btime do_it_many_times($N,$(Ref{Type{Grid{3,7}}}(gd))[])
v1, v2 = Val(3), Val(7)
@btime do_it_many_times2($N,$v1,$v2)

@assert do_it_many_times(N,gd) == do_it_many_times2(N,v1,v2)

```

yields

```julia
124.791 μs (2008 allocations: 31.75 KiB)
124.708 μs (2006 allocations: 31.72 KiB)
18.583 μs (6 allocations: 480 bytes)

```

Unless there’s something wrong with these additional benchmarks, there really does seem to be a performance difference here.

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [July 3, 2024, 3:59pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/9 "2024-07-03T15:59:16Z")

</div>

But here you might run into the no-specialize case, because the `grid` argument is a `DataType`, so no specialization. I.e. `do_it_many_times` must do a dynamic dispatch on `momentum_v1`, I think, resulting in allocations.

---

<div class="post-metadata">

**Author:** ![quantumtwist](https://avatars.discourse-cdn.com/v4/letter/q/ebca7d/32.png) [@quantumtwist](https://discourse.julialang.org/u/quantumtwist)\
**Post date:** [July 3, 2024, 4:03pm UTC](https://discourse.julialang.org/t/bad-performance-from-parametric-struct-dispatch/116594/10 "2024-07-03T16:03:30Z")

</div>

Ahh, I see!

```julia
function do_it_many_times3(N,grid :: T) where {T}
    a = 0
    Random.seed!(4432543)
    for n in 1:N
        config = rand(UInt64) & (1<<22 -1)
        K = momentum_v1(config,grid)
        a += K.n1 #prevent compiler from trivializing loop
    end
    return a
end
@btime do_it_many_times($N,$gd)
@btime do_it_many_times2($N,$v1,$v2)
@btime do_it_many_times3($N,$v1,$v2)

```

yields

```julia
94.584 μs (2006 allocations: 31.72 KiB)
14.333 μs (6 allocations: 480 bytes)
14.667 μs (8 allocations: 512 bytes)

```

Thank you all for helping to resolve this!
