# Poor performance on cluster multithreading

**URL:** <https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248>\
**Category:** Performance\
**Tags:** performance, parallel, multithreading, cluster\
**Created:** [July 8, 2018, 1:09pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248 "2018-07-08T13:09:12Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 8, 2018, 10:07pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/21 "2018-07-08T22:07:15Z")

</div>

That’s why I said you should just allocate the RNG on the thread you are going to use it on. This issue will be automatically taken care of.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 8, 2018, 10:21pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/22 "2018-07-08T22:21:58Z")

</div>

> [@yuyichao](#):
>
> It’s correct to say that the GC percentage is not an accurate representation of how much **ALLOCATION** is hurting you, it is never meant to be and isn’t even called that, but the reason for that is unrelated to what GC does

I see, thank you for the explanation!

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [July 8, 2018, 10:40pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/23 "2018-07-08T22:40:29Z")

</div>

> [@yuyichao](#):
>
> you should just allocate the RNG on the thread you are going to use it on.

Time for an experiment:

```julia
function runme(N=10^8)
    nt = nthreads()
    mt0 = MersenneTwister(1234)
    rngs = randjump(mt0, nt)
    av = zeros(1024,nt)
    @threads for it=1:nt
        # rng = rngs[it] # version 1, global alloc
        rng = copy(rngs[it]) # version 2, threadwise alloc
        a=0.0
        for i=1:N
            a += (rand(rng) - 0.5)
        end
        av[1,it] = a
    end
    av[1,:]
end

```

(Note that the `copy` method is implemented as a deep copy. Striding `av` is pointless here, but helpful if it were used in the hot loop.)

On my system, with 8 threads on 8 cores:  
Version 1: 1.869167 seconds (61 allocations: 122.984 KiB)

Version 2: 0.422015 seconds (101 allocations: 174.734 KiB)

Version 1 is occasionally faster, but I saw good consistency with Version 2.

I am surprised how much this matters, given the large size of the Mersenne Twister state vector.

---

<div class="post-metadata">

**Author:** ![Juser](https://avatars.discourse-cdn.com/v4/letter/j/34f0e0/32.png) [@Juser](https://discourse.julialang.org/u/Juser)\
**Post date:** [July 9, 2018, 1:40am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/24 "2018-07-09T01:40:56Z")

</div>

I’m sure others have mentioned this, but I believe that rand() is not thread-safe and causes problems if multiple threads try to use rand() unless you specify the RNG separately for each thread.

Take a look at this [thread](https://discourse.julialang.org/t/question-about-multi-threading-performance/12075/4) for an example of multi-threaded random number generation.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 9, 2018, 1:46am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/25 "2018-07-09T01:46:32Z")

</div>

> [@Ralph\_Smith](#):
>
> I am surprised how much this matters, given the large size of the Mersenne Twister state vector.

Now you know how much hardware vendors “cheats” and how slow hardwares actually are for serial program accessing memory. Even 1% of \>200x slow down can easily give you a significant global slow down.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 9, 2018, 5:36am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/26 "2018-07-09T05:36:29Z")

</div>

Cool. I updated the comment in the thread I linked, and (besides the obvious advantage of actually being correct / independently random), saw roughly a 50% speedup using `TRNG` to avoid false sharing.  
This was on a Threadripper, which may have some NUMA issues (two separate dies in one physical processor; latency is higher between them[1]). But the end result was about a 20x improvement over single threaded, which seems really good for 16 physical (32 logical) cores.

Oddly, I see a significant slowdown from hoisting the getindex out of the loop:

```julia

function tf!(x::Vector{Matrix{Float64}},N::Int64)
    tmap!(x, x) do x
        id = Threads.threadid()
        n = Threads.nthreads()
        @inbounds for i ∈ 1+(id-1)*N÷n:id*N÷n, j ∈ eachindex(x)
            x[j] += randn(TRNG[Threads.threadid()])
        end
        x
    end
    nothing
end

```

vs

```julia

function tfc!(x::Vector{Matrix{Float64}},N::Int64)
    tmap!(x, x) do x
        id = Threads.threadid()
        n = Threads.nthreads()
        # rng = TRNG[Threads.threadid()]
        rng = copy(TRNG[Threads.threadid()])
        @inbounds for i ∈ 1+(id-1)*N÷n:id*N÷n, j ∈ eachindex(x)
            x[j] += randn(rng)
        end
        x
    end
    nothing
end

```

while both the version that relies on copy, and the TRNG are about the same fast.  
Both are much faster than relying on a regular `randjump`.

[1] [NUMA NUMA: Infinity Fabric Bandwidths - AMD's Future in Servers: New 7000-Series CPUs Launched and EPYC Analysis](https://www.anandtech.com/show/11551/amds-future-in-servers-new-7000-series-cpus-launched-and-epyc-analysis/2)

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [July 9, 2018, 8:27am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/27 "2018-07-09T08:27:18Z")

</div>

Coming rather late to the party here, a small plea. when discussing multi-threaded applications, please specify if you are using all real cores, or hyperthreaded cores.  
Random number discussions aside, I think @DavidBerghaus nails it with the commen tthat the peak of perfomance is at 24 cores.  
I know I am traching my grandmother to suck eggs here, and I actually would appreciate this point being corrected. Hyperthreading to me will only be of benefit in a workload where one computation stalls, so the other computation assigned to that physical core can be ‘switched in’ rapidly. In the past many HPC type codes were known not to benefit from Hyperthreading.

And a second aside, you can test the effects of hyperthreading being switched on or off via a few methods.

a) disable it in the BIOS. OK, that is going to require a reboot and also you may not be able to do that on a remote cluster not under your control

b) In Linux you can disable every odd-numbered processor: echo 0 \> /sys/devices/system/cpu/cpuN/online

c) On Linux use numactl to ‘pin’ processes to particular CPU cores

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 9, 2018, 9:08am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/28 "2018-07-09T09:08:20Z")

</div>

> [@Elrod](#):
>
> This was on a Threadripper, which may have some NUMA issues (two separate dies in one physical processor; latency is higher between them[1]). But the end result was about a 20x improvement over single threaded, which seems really good for 16 physical (32 logical) cores.

Interesting, the situation I described above happened on a server with 2 epyc 32 core cpus. That could be the reason why I saw such a drastic improvement when optimizing for allocations, rather than just the 5% as would be expected by just getting the gc out of the way.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [July 9, 2018, 9:10am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/29 "2018-07-09T09:10:49Z")

</div>

On the hyperthreading, if say I have 128 threads, 64 cores, but tell julia to only use 64 threads, will it automatically distribute the threads over physical cores?

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 9, 2018, 10:54am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/30 "2018-07-09T10:54:50Z")

</div>

> [@louisponet](#):
>
> On the hyperthreading, if say I have 128 threads, 64 cores, but tell julia to only use 64 threads, will it automatically distribute the threads over physical cores?

Good question, I was wondering about that too

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [July 10, 2018, 8:56am UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/31 "2018-07-10T08:56:21Z")

</div>

Regarding the number of cores, there is an issue open at the moment regarding Sys.CPU\_CORES  
[https://github.com/JuliaLang/julia/issues/13901](https://github.com/JuliaLang/julia/issues/13901)

I think Sys.CPU\_LOGICAL\_CORES is a good idea and also Sys.CPU\_PHYSICAL\_CORES  
Maintaining Sys.CPU\_CORES as an equal value to Sys.CPU\_LOGICAL\_CORES is probably justified, but I Can see that leading to Dooooh! moments in the future as people do not appreciate that CPU\_CORES is the logical not physical number.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 10, 2018, 7:02pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/32 "2018-07-10T19:02:46Z")

</div>

I saw your reply here:  
[https://github.com/JuliaLang/julia/issues/17395#issuecomment-232343762](https://github.com/JuliaLang/julia/issues/17395#issuecomment-232343762)

I am using the default sin / cos functions as well, so this might be the issue.  
However, I cannot use your workaround with  
`sin1(x::Float64) = ccall(:sin, Float64, (Float64,), x)`  
since I have to work with BigFloats.  
When I change Float64 to BigFloat, Julia just crashes.

Do you know another way to fix this?

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 10, 2018, 7:03pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/33 "2018-07-10T19:03:51Z")

</div>

> [@DavidBerghaus](#):
>
> since I have to work with BigFloats.

No you are not seeing the same issue.

---

<div class="post-metadata">

**Author:** ![DavidBerghaus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidberghaus/32/4623_2.png) [@DavidBerghaus](https://discourse.julialang.org/u/DavidBerghaus)\
**Post date:** [July 10, 2018, 9:40pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/34 "2018-07-10T21:40:11Z")

</div>

I think the sin/cos are actually the issue for my main algorithm not to perform properly on multithreading.  
I excluded them from the algorithm and finally got the behavior where more threads lead to faster speed.

But you dont know if there is a way for BigFloats?  
I could try to write an algorithm for sin/cos on my own, but they would probably not get the same speed as the default ones…

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [July 11, 2018, 1:45pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/35 "2018-07-11T13:45:38Z")

</div>

I somehow don’t understand all of the issue. We have two potential problems: (1) true sharing: Thread A writes to addr1, thread B reads from addr2, and both are on the same cache line. Yeah, that’s bad and will definitely kill performance. (2) False sharing: The processor only looks at a couple of bits of the address in order to decide where to put it in the cache. These can collide.

Now, what happens for the rng? Every thread needs to grab its mersenne twister. This is a mutable and a half-line (32 bytes); its `mt.idx` will need to get read and written to a lot. If the mersenne twister-structs are on adjacent cache lines, we will get into trouble. Next, we have the buffer (`mt.vals`) and the state (`mt.state`); we will need to access both of them a lot. These are vectors; so we need to hit the vector-struct and the actual buffer. So if any of these guys collide, we are unhappy.

Let’s look at the addresses:

```julia
julia> function fct()
       mt0 = MersenneTwister(0)
       rngs = randjump(mt0, 10)
       end
julia> rngs=fct();
julia> pointer_from_objref.(rngs)
10-element Array{Ptr{Void},1}:
 Ptr{Void} @0x00007f1d0b247dc0
 Ptr{Void} @0x00007f1d0b247e50
 Ptr{Void} @0x00007f1d0b247eb0
 Ptr{Void} @0x00007f1d0b247ee0
 Ptr{Void} @0x00007f1d0b247f10
 Ptr{Void} @0x00007f1d0b247f40
 Ptr{Void} @0x00007f1d0b247fa0
 Ptr{Void} @0x00007f1d0b247fd0
 Ptr{Void} @0x00007f1d0b264010
 Ptr{Void} @0x00007f1d0b264070

```

Urgh. That doesn’t look nice. For example, the counter (`mt.idx`) and buffer (`mt.val`) of the following will share a cache line:

```julia
 Ptr{Void} @0x00007f1d0b247eb0
 Ptr{Void} @0x00007f1d0b247ee0

```

So, in summary, `randjump` should be rewritten or `MersenneTwister` should ~~be padded to 64 bytes and allocated via `posix_memalign`~~ use a ccall of some sort for creation? (nothing should share a cache line with a global counter like `MersenneTwister`, ever; and, somehow even 64-byte length mutable structs can get split over multiple cache lines?)

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 11, 2018, 1:59pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/36 "2018-07-11T13:59:58Z")

</div>

> [@foobar\_lv2](#):
>
> So, in summary, `randjump` should be rewritten or `MersenneTwister` should be padded to 64 bytes and allocated via `posix_memalign` ? (nothing should share a cache line with a global counter like `MersenneTwister` , ever; and, somehow even 64-byte length mutable structs can get split over multiple cache lines?)

None of above. That’s why I repeatedly say that you should allocate **on each thread**. That already guarantee different thread won’t be in the same cache.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [July 11, 2018, 2:06pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/37 "2018-07-11T14:06:53Z")

</div>

So the general solution for thread-local mutable state would be

```julia
using Compat
using Base.Threads
function rng_init()
rngs = Vector{MersenneTwister}(undef, nthreads())
mt0=MersenneTwister(0)
@threads for i=1:nthreads()
rng=randjump(mt0,1)
rngs[i] = rng[1]
end
return rngs
end

```

so that the rngs can be re-used?

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 11, 2018, 2:08pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/38 "2018-07-11T14:08:24Z")

</div>

> [@foobar\_lv2](#):
>
> So the general solution for thread-local mutable state would be

You can use a global vector. Yes that shouldn’t have false sharing issue (edit: it could in principle have false sharing due to cache line affinity limit but that’s a completely different issue…). Not sure what you mean by “re-used”.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [July 11, 2018, 2:11pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/39 "2018-07-11T14:11:32Z")

</div>

> Not sure what you mean by “re-used”.

The previous solution looked like one would create new `MersenneTwister`s during each split (`@threads`) of program flow.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [July 11, 2018, 2:12pm UTC](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248/40 "2018-07-11T14:12:57Z")

</div>

> [@foobar\_lv2](#):
>
> would create new `MersenneTwister` s during each split ( `@threads` ) of program flow.

OK, sure. I assumed that it was written that way for easy testing but yes, the copying in the loop isn’t the point, allocating on the thread is.

[Previous page](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248.md?page=1)

[Next page](https://discourse.julialang.org/t/poor-performance-on-cluster-multithreading/12248.md?page=3)
