# Embarrassingly parallel multi-threading doesn't scale

**URL:** <https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843>\
**Category:** Performance\
**Tags:** multithreading\
**Created:** [October 15, 2021, 5:12pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843 "2021-10-15T17:12:15Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Arnau](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/arnau/32/29962_2.png) [@Arnau](https://discourse.julialang.org/u/Arnau)\
**Post date:** [October 15, 2021, 5:12pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/1 "2021-10-15T17:12:15Z")

</div>

A simple embarrassingly parallel code, with 4 threads, only halves in speed.  
Here is a super simplified case

```julia
:~$ julia -t4

julia> N = Int(1e8)
100000000

julia> @elapsed for i in 1:N
       rand(100)
       end
46.246816922

julia> @elapsed Threads.@threads for i in 1:N
       rand(100)
       end
24.132724014

julia> Threads.nthreads()
4

```

Speed-up is only ~2! How is this possible!?  
I’ve tried different numbers of threads, different numbers N.  
I’ve tried on Jupyter notebooks and with other embarrassingly parallel tasks, but I keep getting the same thing.  
I’ve also tried @belapsed and @btime to make sure this wasn’t a one time thing.  
I have a Thinkpad T14 with Ubuntu 20.04, AMD Ryzen 7 with 8 cores (16 with multi-threading).

I thought maybe it says it’s using 4 threads, but it’s actually just 2 cores with double threading, but I don’t know how to test that.  
Any ideas?  
Thanks

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [October 15, 2021, 5:16pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/2 "2021-10-15T17:16:28Z")

</div>

We had as similar question recently: [Loosing performance with `Threads.@threads` for loop - #6 by Kevin\_Kleinbeck](https://discourse.julialang.org/t/loosing-performance-with-threads-threads-for-loop/69358/6)

Maybe you oversaturate your machine?

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 15, 2021, 5:34pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/3 "2021-10-15T17:34:59Z")

</div>

The `rand(100)` is probably too fast, and also requires a lot of memory allocations, such that the whole computation is probably bounded by memory accesses. These very simple benchmarks usually do not measure anything meaningful in terms of threading performance. In particular, in real applications, you should probably try to avoid allocations as much as possible, exactly to avoid concurrency on memory access. For example, this is probably already better:

```julia
julia> N = 10^5 # :-) 

julia> @btime for i in 1:$N
           for j in 1:100
               rand()
           end
       end
  31.868 ms (0 allocations: 0 bytes)

julia> @btime Threads.@threads for i in 1:$N
           for j in 1:100
               rand()
           end
       end
  10.256 ms (21 allocations: 1.84 KiB)

```

(and `rand()` is a tricky function to include in these benchmarks, because it may have to access a global shared variable to produce the sequence of numbers).

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [October 15, 2021, 6:01pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/4 "2021-10-15T18:01:57Z")

</div>

Hi Leandro,

does the thread scheduling in Julia automatically consider the number of available cores? Even if so, if you start 10000 thread at the same time, isn’ t there some kind of thread switching overhead to expect?

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 15, 2021, 6:22pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/5 "2021-10-15T18:22:09Z")

</div>

These things are lowered roughly in the same way:

```julia
julia> using BenchmarkTools

julia> N = 10^5
100000

julia> @btime Threads.@threads for i in 1:$N
           for j in 1:100
               rand()
           end
       end
  10.329 ms (21 allocations: 1.86 KiB)

julia> @btime Threads.@threads for i in 1:4
           for k in 1:$N÷4
               for j in 1:100
                   rand()
               end
           end
       end
  10.338 ms (21 allocations: 1.86 KiB)

```

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 15, 2021, 6:27pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/6 "2021-10-15T18:27:18Z")

</div>

> [@goerch](#):
>
> does the thread scheduling in Julia automatically consider the number of available cores?

No, you need to start Julia with the number of threads you want to use in the session: [Multi-Threading · The Julia Language](https://docs.julialang.org/en/v1/manual/multi-threading/#Starting-Julia-with-multiple-threads). There is the option to automatically use the number of available threads in the system (`-t auto`).

---

<div class="post-metadata">

**Author:** ![Arnau](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/arnau/32/29962_2.png) [@Arnau](https://discourse.julialang.org/u/Arnau)\
**Post date:** [October 15, 2021, 6:29pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/7 "2021-10-15T18:29:56Z")

</div>

Thanks for the replies.  
Ok this is helpful, but a bit disappointing since in actual real cases I am using rand() a lot and accessing memory in every loop to save the results.

Maybe I’d be better off with Distributed.jl for embarrassingly parallel tasks, then there’s no memory concurrency

---

<div class="post-metadata">

**Author:** ![Arnau](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/arnau/32/29962_2.png) [@Arnau](https://discourse.julialang.org/u/Arnau)\
**Post date:** [October 15, 2021, 6:32pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/8 "2021-10-15T18:32:31Z")

</div>

I don’t think that was the question.

if you launch julia with 4 threads (julia -t4), will it be on 4 different cores (assuming your machine has 4 or more cores), or might it be on two cores double-threading?  
That would explain why performance is only halved

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 15, 2021, 6:36pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/9 "2021-10-15T18:36:38Z")

</div>

> [@Arnau](#):
>
> Ok this is helpful, but a bit disappointing since in actual real cases I am using rand() a lot and accessing memory in every loop to save the results.

To have good scaling the work done on each thread must be enough to compensate the overheads associated to threading. Concerning the memory accesses, you can allocate separate arrays for each thread to store results. If you have an actual application code that you want to scale better, maybe you can ask for help for that specifically, the optimal solutions is certainly dependent on the calculations involved.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 15, 2021, 9:13pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/10 "2021-10-15T21:13:53Z")

</div>

First of all **don’t benchmark loops in the global scope**. It’s really bad for Julia’s performance model. Now, regarding the runtime performance, a lot of this is because of allocation heavy code. You should be trying to reduce allocations in your multithreaded functions by reusing buffers, and if you can’t you should consider using multiprocessing instead.

```julia
julia> f(N) = for i in 1:N
           rand(100)
       end
f (generic function with 1 method)

julia> g(N) = Threads.@threads for i in 1:N
           rand(100)
       end
g (generic function with 1 method)

julia> using Distributed; addprocs(3)
3-element Vector{Int64}:
 2
 3
 4

julia> h(N) = @sync @distributed for i in 1:N
           rand(100)
       end
h (generic function with 1 method)

```

```julia
 N = Int(1e8)

julia> @elapsed f(N)
23.049963313

julia> @elapsed g(N)
18.569370371

julia> @elapsed h(N)
9.839405454

```

If I understand correctly, every time the GC runs, it has to pause all the threads currently which is slowing down the multithreaded code a lot. If you show us something more realistic, it’s possible we can help you eliminate your allocations.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 15, 2021, 9:22pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/11 "2021-10-15T21:22:34Z")

</div>

For example, here’s how I’d cut out the allocations from your example:

```julia
julia> using Random: rand!

julia> function i(N)
           buffers = [Vector{Float64}(undef, 100) for _ ∈ 1:Threads.nthreads()]
           Threads.@threads for i in 1:N
               rand!(buffers[Threads.threadid()])
           end
       end
i (generic function with 1 method)

julia> @belapsed i(N)
3.908675333

```

Here, I preallocate one vector per thread, and then inside the loop, I call `rand!` on that preallocated array, mutating it in place with random numbers instead of creating a new array.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [October 15, 2021, 9:49pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/12 "2021-10-15T21:49:07Z")

</div>

I’d recommend avoiding `buffers[Threads.threadid()]` pattern since it’s very hard to know when it’s OK to use it: [FAQ · FLoops](https://juliafolds.github.io/FLoops.jl/dev/explanation/faq/#faq-state-threadid)

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 16, 2021, 12:19am UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/13 "2021-10-16T00:19:11Z")

</div>

I actually thought of that as I wrote it, but didn’t have ti to chase down the reference and recommended alternative. Thanks for the link.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [October 16, 2021, 4:27am UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/14 "2021-10-16T04:27:08Z")

</div>

Shouldn’t this example be fine though, at least if using `Threads.@threads :static`?  
Because `buffers` was allocated inside `i`, no external tasks will have access to it.  
I’m not sure if the static scheduler really guarantees no task migration, though.

Better would probably be to just `@spawn` a task per thread, and have these tasks reuse a local buffer that they initialize themselves. Each of these could also then explicitly manage their local RNG if necessary/desired.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [October 16, 2021, 8:12am UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/15 "2021-10-16T08:12:15Z")

</div>

> [@Elrod](#):
>
> I’m not sure if the static scheduler really guarantees no task migration, though.

AFAIK, it does. (@vchuravy)

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [October 16, 2021, 9:15am UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/16 "2021-10-16T09:15:19Z")

</div>

`@threads` does currently set the tasks to sticky, but would be good to know if this is an implementation detail or semantically guaranteed. The doc string didn’t mention it.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [October 16, 2021, 9:16am UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/17 "2021-10-16T09:16:54Z")

</div>

My understanding is that it is an implementation detail for `@threads` but guaranteed for `@threads :static`. But I agree, it would be great to make this official in the docs / doc string.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [October 16, 2021, 10:09pm UTC](https://discourse.julialang.org/t/embarrassingly-parallel-multi-threading-doesnt-scale/69843/18 "2021-10-16T22:09:12Z")

</div>

> [@Elrod](#):
>
> Shouldn’t this example be fine though, at least if using `Threads.@threads :static` ?

Yeah, I agree. But the code posted in Discourse will unlikely be used as-is. It’ll be modified “a little bit” by people with various levels of knowledge. It’s possible that people swap `Threads.@threads for` with something else (e.g., direct `@spawn`) and put (say) some `print`-based debug statements.

So, I think recommending `buffers[threadid()]` pattern needs an explanation of why it’s OK here and when it can break. But it’s tedious. So, for simplicity, by default, I prefer to discourage this and recommend a simple generic way. It’s useful sometimes ATM but I think it’s a very implementation-dependent and expert-only pattern.
