# Please help: Poor performance, what am I missing?

**URL:** <https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221>\
**Category:** General Usage\
**Tags:** question, performance\
**Created:** [October 25, 2022, 12:03am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221 "2022-10-25T00:03:54Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jocklawrie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jocklawrie/32/33616_2.png) [@jocklawrie](https://discourse.julialang.org/u/jocklawrie)\
**Post date:** [October 25, 2022, 12:03am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/1 "2022-10-25T00:03:54Z")

</div>

Hello clever people,

I’m having trouble dissecting a performance problem.

MWE below. The computation in the 3rd benchmark is the composition of the computations in the first 2 benchmarks. I expect the 3rd benchmark to take about the total time taken by the first 2 benchmarks. But it’s 2x as long and there’s an unexpected allocation. Clearly I’m missing something fundamental. What’s happening here?

```julia

using BenchmarkTools

"Sum the numeric values, count the non-numeric values"
function processcollection(t)
    total = 0.0
    ncat = 0
    for x in t
        total, ncat = processvalue(x, total, ncat)
    end
    total, ncat
end

processvalue(x::Real, total, ncat) = total + x, ncat
processvalue(x, total, ncat) = total, ncat + 1

d = Dict("a" => (1,2), "b" => ("a", "b", "c"), "c" => (1.1, 2.2, 3.3, 4.4, 5.5, 6.6))

k = "c"
v = d[k]
@benchmark $d[$k] # 20ns, 0 bytes
@benchmark processcollection($v) # 2ns, 0 bytes
@benchmark processcollection($d[$k]) # 45ns, 32 bytes, 1 alloc

```

---

<div class="post-metadata">

**Author:** ![adienes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adienes/32/37459_2.png) [@adienes](https://discourse.julialang.org/u/adienes)\
**Post date:** [October 25, 2022, 12:23am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/2 "2022-10-25T00:23:56Z")

</div>

If you move the lvalues into a function

```julia
function main()
    d = Dict("a" => (1,2), "b" => ("a", "b", "c"), "c" => (1.1, 2.2, 3.3, 4.4, 5.5, 6.6))
    k = "c"
    @benchmark processcollection($d[$k])
end
main()

```

I get `27ns`

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 25, 2022, 12:36am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/3 "2022-10-25T00:36:55Z")

</div>

> [@jocklawrie](#):
>
> `@benchmark processcollection($d[$k])`

In general this will still perform a dynamic dispatch — even if the types of `d` and `k` are known by the compiler (because you interpolated them), it doesn’t know the type of `d[k]` (because your dictionary is heterogeneous).

---

<div class="post-metadata">

**Author:** ![jocklawrie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jocklawrie/32/33616_2.png) [@jocklawrie](https://discourse.julialang.org/u/jocklawrie)\
**Post date:** [October 25, 2022, 1:24am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/4 "2022-10-25T01:24:23Z")

</div>

Doesn’t seem to help - I’m still getting 45ns on my machine.

---

<div class="post-metadata">

**Author:** ![jocklawrie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jocklawrie/32/33616_2.png) [@jocklawrie](https://discourse.julialang.org/u/jocklawrie)\
**Post date:** [October 25, 2022, 1:26am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/5 "2022-10-25T01:26:57Z")

</div>

Thanks Steven.  
I still don’t get why the 3rd benchmark is slower than the sum of the first two. The dynamic dispatch should happen in both the 1st and the 3rd benchmarks right?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 25, 2022, 1:36am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/6 "2022-10-25T01:36:05Z")

</div>

> [@jocklawrie](#):
>
> I still don’t get why the 3rd benchmark is slower than the sum of the first two. The dynamic dispatch should happen in both the 1st and the 3rd benchmarks right?

No, the dynamic dispatch is determining (at runtime) which compiled method of `processcollection` to call based on the type of `d[k]`, and this only happens in the third benchmark.

---

<div class="post-metadata">

**Author:** ![jocklawrie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jocklawrie/32/33616_2.png) [@jocklawrie](https://discourse.julialang.org/u/jocklawrie)\
**Post date:** [October 25, 2022, 2:01am UTC](https://discourse.julialang.org/t/please-help-poor-performance-what-am-i-missing/89221/7 "2022-10-25T02:01:26Z")

</div>

Ah ok, so we’re essentially talking about the distinction between  
`@benchmark processcollection($v) ` and  
`@benchmark processcollection(v) `, which use compile-time dispatch and run-time dispatch respectively, and the latter is the same as the 3rd benchmark.

Thanks again, most helpful.  
Jock
