# Parallelizaton on GPU slower than on CPU...?

**URL:** <https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587>\
**Category:** Performance\
**Tags:** gpu\
**Created:** [January 20, 2020, 2:17pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587 "2020-01-20T14:17:08Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [January 20, 2020, 2:17pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/1 "2020-01-20T14:17:08Z")

</div>

I built a new PC and I’ve been toying around with some code just to see how it performs and I’m trying to understand why the following code (from the CUDA.jl docs) runs slower on my GPU than on my CPU (I have a Ryzen 9 3950x CPU (16 core/32 thread) and an RTX 2080 Super GPU):

```julia
using BenchmarkTools
using CuArrays
using Test

N = 2^21
x = fill(1.0f0, N) # a vector filled with 1.0 (Float32)
y = fill(2.0f0, N) # a vector filled with 2.0
y .+= x   

function sequential_add!(y, x)
    for i in eachindex(y, x)
        @inbounds y[i] += x[i]
    end
    return nothing
end

fill!(y, 2)
sequential_add!(y, x)
@test all(y .== 3.0f0)

function parallel_add!(y, x)
    Threads.@threads for i in eachindex(y, x)
        @inbounds y[i] += x[i]
    end
    return nothing
end

fill!(y, 2)
parallel_add!(y, x)
@test all(y .== 3.0f0)

# Parallelizaton on the GPU
x_d = CuArrays.fill(1.0f0, N) # a vector stored on the GPU filled with 1.0 (Float32)
y_d = CuArrays.fill(2.0f0, N) # a vector stored on the GPU filled with 2.0
y_d .+= x_d
@test all(Array(y_d) .== 3.0f0)

function add_broadcast!(y, x)
    CuArrays.@sync y .+= x
    return
end

```

The results are here:

```julia
julia> @btime sequential_add!($y, $x)
  254.301 μs (0 allocations: 0 bytes)

julia> @btime parallel_add!($y, $x)
  44.499 μs (114 allocations: 13.67 KiB)

julia> @btime add_broadcast!($y_d, $x_d)
  106.000 μs (56 allocations: 2.22 KiB)

```

As you can see, the CPU crushes the GPU with this computation (I love my new CPU 🥰)

Lastly, for a 16 core, 32 thread CPU, is it okay to set `JULIA_NUM_THREADS` to 32, or should it equal the number of physical cores? I currently have it set at 16.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [January 20, 2020, 3:43pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/2 "2020-01-20T15:43:46Z")

</div>

Are the results the same if you further increase the length of the vectors?

> [@mthelm85](#):
>
> Lastly, for a 16 core, 32 thread CPU, is it okay to set `JULIA_NUM_THREADS` to 32, or should it equal the number of physical cores? I currently have it set at 16.

I would keep it at 16, the hyper threads do not have their own cache memory and would mostly compete for resources with the native threads.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [January 20, 2020, 6:13pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/3 "2020-01-20T18:13:53Z")

</div>

> Are the results the same if you further increase the length of the vectors?

I tried with `N = 2^20`, `N = 2^21`, `N = 2^22` and `N = 2^23` and `parallel_add!` was faster than `add_broadcast!`. However, at `N = 2^27`, `add_broadcast!` is _much_ faster:

```julia
julia> @btime sequential_add!($y, $x)
  60.774 ms (0 allocations: 0 bytes)

julia> @btime parallel_add!($y, $x)
  57.521 ms (114 allocations: 13.67 KiB)

julia> @btime add_broadcast!($y_d, $x_d)
  3.745 ms (56 allocations: 2.22 KiB)

```

So, to decide whether or not it’s worth doing something on the GPU, is the best way trial-and-error, or is there some sort of rule of thumb to go by?

Thanks 😀

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [January 20, 2020, 6:38pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/4 "2020-01-20T18:38:22Z")

</div>

I can imagine that it varies a lot with _what_ you do with the array. Try `exp` and I bet that the GPU will be faster much earlier.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 21, 2020, 12:54am UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/5 "2020-01-21T00:54:44Z")

</div>

Those are some impressive numbers on the 1950X!

Note that even though the two arrays only take up 16 MiB (2^21 \* 4 \* 2 / 2^20 = 2^4), the computation is memory bound.

```julia
julia> N = 2^21
2097152

julia> flops = 10^6 * N / 44.499
4.712807029371446e10

```

I don’t know what clock speed your CPU runs at all-core, so I’ll pick 4 GHz:

```julia
julia> Hz = 4e9; fma_per_clock = 2; flop_per_fma = 16; cores = 16;

julia> Hz * fma_per_clock * flop_per_fma * cores
2.048e12

julia> ans / flops
43.4560546875

```

Your CPU was mostly sitting, waiting for data. For every nanosecond it spent computing, there were 40 doing nothing.

For comparison, on my 10980XE, my sequential and parallel times were 705 and 58 microseconds.  
Thus, my numbers are

```julia
julia> Hz = 4.1e9; fma_per_clock = 2; flop_per_fma = 32; cores = 18;

julia> Hz * fma_per_clock * flop_per_fma * cores
4.7232e12

julia> ans / (10^6 * N / 58)
130.62744140625

```

Yikes. My ratio was about 130.

I don’t know much about GPU computing, but I bet you couldn’t bring it’s number crunching power to bear. Longer vectors would just make the memory problems worse.

I also don’t enough yet about memory to say anything about TLB misses vs memory bandwidth, but I’ll start looking into that sort of thing one day.

For memory bound operations, memory performance dominates. Regardless of the reason, the Ryzen 3950X looks amazing here.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [January 21, 2020, 2:04am UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/6 "2020-01-21T02:04:25Z")

</div>

> For memory bound operations, memory performance dominates.

I kept this in mind for my build and went with DDR4-3600 RAM as well as an M.2-2280 NVME SSD 😁. I also got inspired to start [a thread for showing off Julia performance on PCs that people build/have](https://discourse.julialang.org/t/show-off-julia-performance-on-your-pc/33606) so check it out! Thanks so much for your response!!

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 21, 2020, 1:49pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/7 "2020-01-21T13:49:04Z")

</div>

But because both arrays here are still only 16 MiB, those don’t matter for this specific benchmark.  
The 10980XE has 18MiB of shared L3 cache, while the 1950X has 16 MiB of L3 cache per 4 cores (for 64 MiB total).

If I recall correctly, on Zen3, much of the on CPU memory’s clock matches the RAM clock up until 3600 MHz. Meaning you may have those speeds set much higher than I do.  
I didn’t overclock/adjust “uncore” performance at all, and have no idea how much ground I can gain from that. Probably worth at least looking at, but whatever I do it’ll probably be mild since I don’t really want to risk crashing the computer. I wouldn’t expect to gain much ground on your benchmark performance here.

FWIW, I’m also on an M.2 SSD, but I don’t remember which at the moment, and DDR4-3200 RAM (14 CAS latency, IIRC).

You should try some more compute-heavy benchmarks in that thread. I’d add matmul, at least ;).

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [January 21, 2020, 2:32pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/8 "2020-01-21T14:32:39Z")

</div>

> [@Elrod](#):
>
> I didn’t overclock/adjust “uncore” performance at all, and have no idea how much ground I can gain from that. Probably worth at least looking at

From what I gather reading discussions about this, the concept of a clock frequency is pretty fluid in late Ryzens, and the CPU monitors itself to keep its performance close to optimal. It is of course possible that one can improve on the defaults, but the gains seems to be rather small.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 21, 2020, 2:43pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/9 "2020-01-21T14:43:50Z")

</div>

[Uncore](https://en.wikipedia.org/wiki/Uncore) refers to:

> Uncore functions include [QPI](https://en.wikipedia.org/wiki/Intel_QuickPath_Interconnect) controllers, [L3 cache](https://en.wikipedia.org/wiki/L3_cache), [snoop agent](https://en.wikipedia.org/wiki/Memory_coherence) [pipeline](https://en.wikipedia.org/wiki/Instruction_pipeline), on-die [memory controller](https://en.wikipedia.org/wiki/Memory_controller), and [Thunderbolt controller](https://en.wikipedia.org/wiki/Thunderbolt_(interface))

This (L3 cache) is probably the most important (hardware) capability being benchmarked here.

While the core includes the execution units as well as the L1 and L2 cache.

Apparently “uncore” is an Intel term, so AMD may do things differently.  
I believe that the infinity fabric clock matches the RAM clock rate up to 3600 MHz (which is an overclock; the “maximum” is 3200). After this the infinity fabric slows down to 2-to-1, making 3600 “optimal”.  
However, I don’t know what all infinity fabric entails – whether it includes L3 cache performance.

I also much prefer’s AMD’s intelligent clock-speed algorithms.

---

<div class="post-metadata">

**Author:** ![randyzwitch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/randyzwitch/32/15955_2.png) [@randyzwitch](https://discourse.julialang.org/u/randyzwitch)\
**Post date:** [January 21, 2020, 2:46pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/10 "2020-01-21T14:46:32Z")

</div>

> [@mthelm85](#):
>
> So, to decide whether or not it’s worth doing something on the GPU, is the best way trial-and-error, or is there some sort of rule of thumb to go by?

My post is a bit stale at this point, but I did an example to understand this same question:

> **[randyzwitch.com | Parallelizing Distance Calculations Using A GPU With...](https://randyzwitch.com/cudanative-jl-julia/)**
>
> Using CUDAnative.jl, you can access the power of GPU parallelization while still writing high-level Julia code. It's not unreasonable to get speedups of 20x or more through GPU parallelization.

What (I believe) I demonstrated was that for smaller problems, the data transfer time eats away at the potential speedup for GPU (as seen by the horizontal line up to 1000x1000 matrix). Once you problem gets larger, then you start to see the GPU start to shine.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [January 21, 2020, 2:54pm UTC](https://discourse.julialang.org/t/parallelizaton-on-gpu-slower-than-on-cpu/33587/11 "2020-01-21T14:54:51Z")

</div>

@randyzwitch Really nice post, thank you.
