# Poor performance on cluster multithreading

**URL:** <https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248>\
**Category:** Performance\
**Tags:** performance, parallel, multithreading, cluster\
**Created:** [July 8, 2018, 1:09pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248 "2018-07-08T13:09:12Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 1:09pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/1 "2018-07-08T13:09:12Z")

</div>

I am currently running Julia on a computing cluster with 56 threads.  
To test the multithreading capacity, I am generating random numbers:

function test()  
tic()  
Threads.@threads for i=1:100000  
rand(1000000)  
end  
toc()  
end

and measure the computing time for different amounts of threads (changed through “export JULIA\_NUM\_THREADS=”):

#Threads Computing time in s  
1 234.69  
4 82.99  
14 38.15  
16 35.85  
28 35.05  
40 39.25  
56 87.07

Does anyone have an idea what could cause this behavior?

Thanks in advance

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [July 8, 2018, 3:30pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/2 "2018-07-08T15:30:11Z")

</div>

> [@DavidBerghaus](#):
>
> Does anyone have an idea what could cause this behavior?

You’re using a global RNG object. Instead, make local RNGs and pass them to `rand`. Maybe that helps?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 8, 2018, 3:53pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/3 "2018-07-08T15:53:50Z")

</div>

Additionally using `rand` without local RNG will produce wrong results. See last part of section [https://docs.julialang.org/en/latest/manual/parallel-computing/#The-@threads-Macro-1](https://docs.julialang.org/en/latest/manual/parallel-computing/#The-@threads-Macro-1) in the Julia manual (on Julia 0.6.3 the example should read a bit differently as `randjump` produces a vector, but the idea is the same - random number generation is not thread safe as `rand` does not perform locking of global random number generator).

The slowdown - in general - is probably due to [false sharing](https://en.wikipedia.org/wiki/False_sharing) as each thread invalidates cache of other threads. But this is a secondary issue - the bigger problem is that simply this code will not produce correct results if I am not missing something. The additional reason can be that most probably your cluster has 56 virtual cores and only 23 physical cores and using two virtual cores on one physical core can lead to performance degradation.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 4:37pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/4 "2018-07-08T16:37:12Z")

</div>

You mean for example adding srand(i) inside the loop to create local RNGs? I tried that and still experienced the same behavior.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 4:37pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/5 "2018-07-08T16:37:13Z")

</div>

You are right, but this function is not actually part of the project that I am working on, so I wasnt thinking about the results.  
I just wanted to create a simple example that should be trivially multi threadable.  
So I created a small function with random numbers and noticed that it has the same performance behavior as the main algorithm that I am using, where lots of threads lead to bad performance.

Do you have an idea why the false sharing could appear? Because from my understanding, the threads should be able to perform completely separate.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 4:37pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/6 "2018-07-08T16:37:15Z")

</div>

Reply to your updated post:  
I had that idea as well after I made this post, but it has 28 physical cores and performes best at around 24.  
For the “real” algorithm that I am using it actually has the fastest speed with only 4 threads, which confused me a lot.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [July 8, 2018, 4:43pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/7 "2018-07-08T16:43:59Z")

</div>

> [@DavidBerghaus](#):
>
> You mean for example adding srand(i) inside the loop to create local RNGs? I tried that and still experienced the same behavior.

`srand(i)` does not create local RNGs. That just seeds the global RNG.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 8, 2018, 4:58pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/8 "2018-07-08T16:58:01Z")

</div>

How many allocations happen in the real algorithm?

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 5:25pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/9 "2018-07-08T17:25:09Z")

</div>

Thats hard to tell since I am calculating an irrational number numerically, so I never really get to an end and the amount of digits increases over time.  
I let the algorithm run for 32s and got the following results with the @time function:  
32.894087 seconds (140.62 M allocations: 3.837 GiB, 5.24% gc time)

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 8, 2018, 7:05pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/10 "2018-07-08T19:05:52Z")

</div>

There are two issues:

1. understanding false sharing: the simplest way to get an intuition when this can happen is when two threads modify regions of memory that are close to each other; this does not have to be the same memory address - it is enough that distance in memory location is close enough that it falls into the same “chunk” of memory that is kept in CPU cache; in such a situation when thread 1 changes memory address A then thread 2 that uses memory address B that is close to address A gets its cache invalidated and this causes performance regression;
2. in general your computer has many different resources that can be a performance bottleneck (CPU, memory, disk IO etc.) - it seems that simply other resource than CPU is a bottleneck in your code (assuming the code is correct otherwise 🙂).

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 8, 2018, 7:31pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/11 "2018-07-08T19:31:11Z")

</div>

What I’ve found from experience is that any allocations at all (unless trivial i.e. in the kb range) is absolutely detrimental for performance, because gc is still not multithreaded. So even though you might see 5% it might actually impact things more as far as I understand it. To quickly check whether this is the issue, try checking the scaling with `gc_enable(false)`. If there is an improvement you should try to work with some caches (1 per thread that get allocated before the loop, and are used in calculations during the loop).  
By eliminating all but a couple of kb of allocations I managed to get almost linear scaling up to 64 cores for some FEM calculations.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 8, 2018, 8:25pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/12 "2018-07-08T20:25:34Z")

</div>

> [@bkamins](#):
>
> understanding false sharing

In this case multiple threads are accessing the same block of memory. It’s true sharing instead of false sharing.

> [@louisponet](#):
>
> So even though you might see 5% it might actually impact things more as far as I understand it.

5% means just that, 5% time in GC, you get no more than 5% off if you disable GC. That said, allocation can easily kill peformance for a number of reasons. It’s just unrelated to GC being multithreaded or not.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 8, 2018, 8:37pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/13 "2018-07-08T20:37:20Z")

</div>

Thank you very much for the advise!  
I might have underestimated the impact of allocations on the performance. I will read through the performance optimization guide and try to avoid allocations as much as I can.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 8, 2018, 8:50pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/14 "2018-07-08T20:50:13Z")

</div>

Agreed. Sorry for being imprecise. I meant that even if this is fixed using `randjump` to get separate generators under Julia 0.6 there is a risk of false sharing from my experience.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 8, 2018, 8:58pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/15 "2018-07-08T20:58:01Z")

</div>

> [@bkamins](#):
>
> there is a risk of false sharing from my experience

Correct, and it’s actually measureable. The different RNG’s have to be allocated on each threads instead of allocating on the same thread.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 8, 2018, 9:02pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/16 "2018-07-08T21:02:14Z")

</div>

> [@Question about Multi-threading Performance](https://discourse.julialang.org/t/question-about-multi-threading-performance/12075/3):
>
> randn isn’t threadsafe. using Compat, Compat.Random const twisters = [MersenneTwister() for i ∈ 1:Threads.nthreads()]; #Multithreaded. function tf!(x::Vector{Matrix{Float64}},N::Int64) Threads.@threads for ii=1:N id = Threads.threadid() twister = twisters[id] @inbounds x\_thd = x[id] for nn=1:100 for mm=1:100 @inbounds x\_thd[mm,nn] += randn(twister) end end end return nothing end yields: julia\> @btime…

You can solve the problem like this:

```julia
using Compat, Compat.Random

const twisters = [MersenneTwister() for i ∈ 1:Threads.nthreads()];

#Multithreaded.
function tf!(x::Vector{Matrix{Float64}},N::Int64)
    Threads.@threads for ii=1:N
        id = Threads.threadid()
        twister = twisters[id]
        @inbounds x_thd = x[id]
        for nn=1:100
            for mm=1:100
                @inbounds x_thd[mm,nn] += randn(twister)
            end
        end
    end
    return nothing
end

```

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 8, 2018, 9:15pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/17 "2018-07-08T21:15:42Z")

</div>

> [@Elrod](#):
>
> const twisters = [MersenneTwister() for i ∈ 1:Threads.nthreads()];

This will actually have false sharing. You need to allocate the rng from each threads.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 8, 2018, 9:46pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/18 "2018-07-08T21:46:20Z")

</div>

Actually there are two problems with this code:

1. `MersenneTwister()` does not ensure that PRNGs are non-overlapping (although the risk of overlap is very small - but this is the reason why `randjump` is provided)
2. The simplest you can do is either do `deepcopy` in each thread an element used by this thread of an array of initiated PRNGs in a single thread or use a “separator” object (one PRNG that is discarded to separate data in threads) similarly to what I proposed in [KissThreading.jl/KissThreading.jl at master · mohamed82008/KissThreading.jl · GitHub](https://github.com/bkamins/KissThreading.jl/blob/master/src/KissThreading.jl#L7). The discarding approach might not be best if you have [NUMA](https://en.wikipedia.org/wiki/Non-uniform_memory_access) issues in your infrastructure.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 8, 2018, 9:46pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/19 "2018-07-08T21:46:38Z")

</div>

I see, then I was mistaken. For some reason though, I also remember that in my case I had allocations supposedly taking up around 5-10%, but reducing them yielded a huge speedup. It felt as if when the gc hit, it would go through all the small chunks that each of the 64threads allocated separately (going through the small chunks that were allocated each loop iteration rather than doing each thread’s total allocated chunk), thus taking a very long time. Not sure if that was exactly what was going on though, I might be completely wrong.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 8, 2018, 10:06pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/20 "2018-07-08T22:06:07Z")

</div>

> [@louisponet](#):
>
> It felt as if when the gc hit, it would go through all the small chunks that each of the 64threads allocated separately (going through the small chunks that were allocated each loop iteration rather than doing each thread’s total allocated chunk), thus taking a very long time.

This is correct. However, that’s exactly what takes 5% of the time!

> [@louisponet](#):
>
> I also remember that in my case I had allocations supposedly taking up around 5-10%, but reducing them yielded a huge speedup.

Reducing allocation can certainly speed things up by a lot. However, at most 5% of those will come from the GC in this case. It’s correct to say that the GC percentage is not an accurate representation of how much **ALLOCATION** is hurting you, it is never meant to be and isn’t even called that, but the reason for that is unrelated to what GC does.

[Next page](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248.md?page=2)
