# Using and understanding multi-threading

**URL:** https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682
**Category:** Performance
**Tags:** multithreading
**Created:** [January 11, 2024, 1:10pm UTC](https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682 "2024-01-11T13:10:14Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)
#### Post date: [January 11, 2024, 1:10pm UTC](https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682/1 "2024-01-11T13:10:14Z")

</div>

so, I am trying to understand how multi threading works, I started julia as `julia -t 16` on a 32 cores node where I reserved 8, then I run the example from the documentation [here](https://docs.julialang.org/en/v1/manual/multi-threading/#Using-@threads-without-data-races) these are the lines

```julia
function sum_single(a)
     s = 0
     for i in a
         s += i
     end
     s
 end

function sum_multi_good(a)
     chunks = Iterators.partition(a, length(a) ÷ Threads.nthreads())
     tasks = map(chunks) do chunk
         Threads.@spawn sum_single(chunk)
     end
     chunk_sums = fetch.(tasks)
     return sum_single(chunk_sums)
 end

```

then I did the benchmarks (after loading `Benchmarktools`) and here are the results:

```julia
julia> @benchmark sum_single(1:1_000_000)
BenchmarkTools.Trial: 10000 samples with 1000 evaluations.
 Range (min … max): 3.246 ns … 13.705 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 3.250 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 3.271 ns ± 0.252 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

  █▇ ▂ ▁ ▁
  ██▄▃▁▅▇▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▄▄▅▁██▇▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ █
  3.25 ns Histogram: log(frequency) by time 3.5 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

```julia
@benchmark sum_multi_good(1:1_000_000)
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 11.700 μs … 4.115 ms ┊ GC (min … max): 0.00% … 96.91%
 Time (median): 34.866 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 38.832 μs ± 57.699 μs ┊ GC (mean ± σ): 1.98% ± 1.37%

                   ▂▅▇█▅▃▁
  ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂▃▆████████▇▆▇▇▆▆▆▅▅▅▅▅▄▄▄▄▃▄▄▅▆▅▅▅▅▄▄▃▃▃▂▂▂▂ ▃
  11.7 μs Histogram: frequency by time 62.3 μs <

 Memory estimate: 8.89 KiB, allocs estimate: 106.

```

why is the parallel version so much slower than the single threaded? (11.7e3/3.25)!

In general I don’t seem to get significant gains , from multithreading, I am using it to evaluate the stiffness matrices of individual elements in a finite element program before assembling the structure stiffness matrix, each element evaluation is quite expensive but they independent from the others, so it seems sort of the best case scenario for parallel computing but still I don’e seem to get significant time savings

many thanks in advance!

---

<div class="post-metadata">

### Author: ![Liozou](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liozou/32/35118_2.png) [@Liozou](https://discourse.julialang.org/u/Liozou)
#### Post date: [January 11, 2024, 2:01pm UTC](https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682/2 "2024-01-11T14:01:30Z")

</div>

The compiler is simply too smart for your simple example!  
To see this, you can have a look at `@code_llvm sum_single(1:1_000_000)` (or `@code_native` if you prefer) to inspect how the compiler has digested your call, and you will notice that there is no loop remaining… That’s because LLVM has built-in optimization for this particular kind of computation, so it will just look at the first and last element of the range, and give you your result. Obviously, doing that computation is much cheaper than allocating tasks to do parallel stuff, so however fast `sum_multi_good` can go, `sum_single` will always be faster. By the way, you can notice that the timing for `sum_single` is around the nanosecond, which is of the order of precision of `@benchmark` (benchmarking `2*7` would probably yield a similar timing), which also confirms that your `sum_single` call is barely doing any computation. You could also check that the timing does not depend on the length of the input range.

To bypass this cleverness of the compiler, you should input something that it cannot optimize away. The easiest is probably to feed `sum_single` with an array, instead of a range. This is what I get using 10 threads:

```julia
julia> @benchmark sum_single($(collect(1:1_000_000)))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 356.141 μs … 1.310 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 363.795 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 368.972 μs ± 18.412 μs ┊ GC (mean ± σ): 0.00% ± 0.00%

   ▁▃▃█▂▁                                                       
  ▄██████▆▅▄▄▄▅▅▄▄▄▄▄▄▄▄▄▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▁▂▂▂▂▂▂ ▃
  356 μs Histogram: frequency by time 431 μs <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sum_multi_good($(collect(1:1_000_000)))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 40.893 μs … 6.037 ms ┊ GC (min … max): 0.00% … 94.62%
 Time (median): 54.895 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 57.975 μs ± 78.316 μs ┊ GC (mean ± σ): 1.81% ± 1.34%

           ▁▃▅▅▇▇▇██▇▅▅▄▃▂▂▁                                   
  ▁▁▂▂▂▃▄▆▇██████████████████▇▇▇▅▅▅▅▄▅▄▄▄▃▄▃▃▃▂▃▂▂▂▂▂▂▂▂▁▂▁▁▁ ▄
  40.9 μs Histogram: frequency by time 83.7 μs <

 Memory estimate: 11.33 KiB, allocs estimate: 130.

```

Just a last note on performance: your current code is type unstable, because `fetch` is inherently type unstable (it’s impossible for the compiler to guess the return type of a `Task`), so you should annotate the call to give the type to the compiler. In this instance, you could write

```julia
T = eltype(a)
chunk_sums = T[fetch(t) for t in tasks]

```

instead of

```julia
chunk_sums = fetch.(tasks)

```

for example
