# Gcc vs Threads.@threads vs Threads.@spawn for large loops

**URL:** <https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273>\
**Category:** Julia at Scale\
**Created:** [February 6, 2020, 8:42pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273 "2020-02-06T20:42:40Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 6, 2020, 8:42pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/1 "2020-02-06T20:42:40Z")

</div>

Hi, Schönauer vector triad again, see also the [benchmarking site of Georg Hager](https://blogs.fau.de/hager/archives/tag/benchmarking).

This time we compare multithreading vs. scalar, and also compare to gcc.

![](https://global.discourse-cdn.com/julialang/original/3X/f/4/f42f5f33362d8d3993afdf6fc1dde62f480cca63.png)

[Here](https://github.com/j-fu/julia-tests/tree/master/parallel) ist the generating code.

We see that for scalar performance, julia shines vs. `gcc -Ofast`. For multithreading performance (4 threads), the picture appears to be different. It seems that the bookkeeping overhead for handling threading is still larger compared to what we see for gcc. For small array sizes we see this overhead problem also for gcc. For large array sizes, performance is limited by memory access (it’s a laptop…).

This is the @threads based loop:

```julia
            Threads.@threads for i=1:N
                @inbounds @fastmath d[i]=a[i]+b[i]*c[i]
            end

```

And this is the gist of the faster, @spawn implementation:

```julia
mapreduce(task->fetch(task),+,[Threads.@spawn _kernel(a,b,c,d,loop_begin[i],loop_end[i]) for i=1:ntasks])

```

A very similar picture one finds in a [2013 post](https://blogs.fau.de/hager/archives/6883) by G. Hager on intel vs. gnu compiler:

 ![](https://global.discourse-cdn.com/julialang/original/3X/7/6/769d8b0182b384c422fcceecd274cb43f9004757.png)

He also suspects barrier performance for gcc to be the reason for the performance difference.

I am aware of the fact that it is still early times for multithreading in Julia, and things are clearly marked as experimental, so this post is meant as an encouragement to continue the endeavour of trying to be on par with C performance-wise (and thanks anyway for starting this!)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [February 6, 2020, 9:22pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/2 "2020-02-06T21:22:06Z")

</div>

Have you looked at [GitHub - tro3/ThreadPools.jl: Improved thread management for background and nonuniform tasks in Julia. Docs at https://tro3.github.io/ThreadPools.jl](https://github.com/tro3/ThreadPools.jl) or [https://github.com/chriselrod/LoopVectorization.jl](https://github.com/chriselrod/LoopVectorization.jl)

---

<div class="post-metadata">

**Author:** ![kpamnany](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kpamnany/32/206168_2.png) [@kpamnany](https://discourse.julialang.org/u/kpamnany)\
**Post date:** [February 6, 2020, 11:19pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/3 "2020-02-06T23:19:17Z")

</div>

Thanks for the post. We do have some work to do in reducing overheads. In particular, `@threads for` is not an optimized implementation.

However, I do want to note that a fairer comparison would be with `#pragma omp parallel for schedule(dynamic)` since Julia’s multi-threading is inherently dynamically scheduled. This is a part of how we achieve nested parallelism which is very difficult to accomplish with OpenMP. A less happy consequence is that we are unlikely to match a statically scheduled OpenMP parallel loop in terms of overheads – we simply do far more.

---

<div class="post-metadata">

**Author:** ![Crown421](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/crown421/32/9984_2.png) [@Crown421](https://discourse.julialang.org/u/Crown421)\
**Post date:** [February 7, 2020, 12:31am UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/4 "2020-02-07T00:31:45Z")

</div>

> [@kpamnany](#):
>
> In particular, `@threads for` is not an optimized implementation

Out of personal interest, is there a recommended alternative? Or this a case of, in general this is thing to use, we just haven’t optimized it yet?

---

<div class="post-metadata">

**Author:** ![kpamnany](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kpamnany/32/206168_2.png) [@kpamnany](https://discourse.julialang.org/u/kpamnany)\
**Post date:** [February 7, 2020, 1:31am UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/5 "2020-02-07T01:31:45Z")

</div>

There is no “recommended” alternative, but there are at least 2 and probably more packages that layer additional/different parallelism functionality above the language’s runtime. We do plan to enhance the parallel loop support in the runtime and it may stay in this form or have a different name/semantics – not concrete yet.

---

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 7, 2020, 11:08am UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/6 "2020-02-07T11:08:51Z")

</div>

Hi, thanks for the hints.

I am trying out ThreadPools, but need to understand more about this package before I would say anything.

`@avx` doesn’t give any additional performance here as it is memory acces bound, but surprisingly it changes the picture with caching and SharedArrays,  
see [my other post](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269) which I will update.

---

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 7, 2020, 1:33pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/7 "2020-02-07T13:33:49Z")

</div>

Thank you for the info. Below I update the figure with the result for  
`#pragma omp for schedule(dynamic, N/omp_max_threads())`. Indeed this is slower than omp with static scheduling. I tried out several different ways of “by hand” scheduling for Julia, but these don’t make much of a difference.  
I think that these large, statically scheduleable loops are not that seldom in PDE numerics, especially iterative solvers. But then, of course, not everything can be done at once…

![](https://global.discourse-cdn.com/julialang/original/3X/f/9/f9ca2e332ad9d9d2f3ae0f8b886b02f81ba5ba6b.png)

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [February 7, 2020, 2:07pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/8 "2020-02-07T14:07:11Z")

</div>

I thought `@threads for` was a static scheduling over the range you give it. But perhaps the underlying scheduler doesn’t exploit this so there is still the same overhead as for dynamic scheduling.

---

<div class="post-metadata">

**Author:** ![kpamnany](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kpamnany/32/206168_2.png) [@kpamnany](https://discourse.julialang.org/u/kpamnany)\
**Post date:** [February 7, 2020, 3:51pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/9 "2020-02-07T15:51:19Z")

</div>

A statically scheduled loop in OpenMP splits the loop range over `num_threads` and typically uses a tree to fan-out the work to the threads in parallel. At the end of the loop, the inverse of the tree is typically used for a barrier. This broadcast-barrier pair of synchronization constructs are pretty much the entire overhead for the loop and these are very well studied – each takes only hundreds to a few thousand cycles, depending on the processor and the number of threads.

But what happens if in the loop body you call another function that also has a parallel loop? With OpenMP you have to analyze your call graph, determine thread allocation at each level, and then carefully set up thread affinities, and possibly environment variables for libraries, etc. in order to use static scheduling all the way down. It is not impossible, just very very hard. So OpenMP added tasks and teams. And those aren’t pervasive in libraries anyway.

What I’m getting at is that it isn’t just variable duration loop iterations that require dynamic scheduling of the sort Julia’s scheduler manages.

As you say, ‘not everything can be done at once’, or IOW, there’s no magic bullet. Nonetheless, we still hope to improve the common case.

---

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 7, 2020, 3:56pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/10 "2020-02-07T15:56:53Z")

</div>

Hi, some update here:  
While the above results have been obtained on a laptop with 1 NUMA node + 12MB L3 cache, here I show the result from a Xeon server with 2 NUMA nodes with 256GB + 40MB L3 cache each. As above, the parallel case uses 4 threads.

![](https://global.discourse-cdn.com/julialang/original/3X/d/7/d74eb8e530bbb9b536fa7abc99c55f053e6ef6b6.png)

So here the picture is a bit different. The processor is slightly slower, and cache is larger, so we see a much larger region where parallelization works. Moreover we have two lanes to memory. Now, `@threads` makes a better impression ;-).  
So if chunk sizes are large enough, Julia appears to be on par.

---

<div class="post-metadata">

**Author:** ![kpamnany](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kpamnany/32/206168_2.png) [@kpamnany](https://discourse.julialang.org/u/kpamnany)\
**Post date:** [February 7, 2020, 4:04pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/11 "2020-02-07T16:04:58Z")

</div>

> I thought `@threads for` was a static scheduling over the range you give it. But perhaps the underlying scheduler doesn’t exploit this so there is still the same overhead as for dynamic scheduling.

It is. But the broadcast+barrier implementations here are completely serial. As `N` increases, these overheads become relatively smaller (edit: as shown in the most recent graph), but there’s lots of scope for improvement. 🙂

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [June 26, 2020, 7:11pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/12 "2020-06-26T19:11:35Z")

</div>

> [@tbeason](#):
>
> Have you looked at [GitHub - tro3/ThreadPools.jl: Improved thread management for background and nonuniform tasks in Julia. Docs at https://tro3.github.io/ThreadPools.jl](https://github.com/tro3/ThreadPools.jl) or [GitHub - JuliaSIMD/LoopVectorization.jl: Macro(s) for vectorizing loops.](https://github.com/chriselrod/LoopVectorization.jl)

What is the difference between those two packages?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [June 26, 2020, 7:21pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/13 "2020-06-26T19:21:29Z")

</div>

One is for threaded execution of the code, the other takes advantage of multiple instructions per cycle when working on array data.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [June 26, 2020, 7:36pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/14 "2020-06-26T19:36:51Z")

</div>

Then ThreadPools.jl is an alternative to _Threads_ .@ _spawn_?  
And is it a good idea to try to use both ThreadPools and LoopVectorization at the same time?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [June 26, 2020, 7:45pm UTC](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/15 "2020-06-26T19:45:03Z")

</div>

I do use threads and LoopVectorization in the same program.  
I believe the ThreadPools free the programmer from having to  
manage the threads manually. I haven’t used it myself.
