# Question for lower performance by using @threads in for loop

**URL:** <https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [July 9, 2021, 1:17am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315 "2021-07-09T01:17:16Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Amro](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@Amro](https://discourse.julialang.org/u/Amro)\
**Post date:** [July 9, 2021, 1:17am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/1 "2021-07-09T01:17:17Z")

</div>

Hello,  
I am surprised that when using the below code, I have the following execution time:

```julia
using .Threads

n=10000;
s=1
Y=zeros(n,n);

@time begin
    for i in 1:n
        for m in 1:30
            if s == 1
                Y[i,i] += m;
            end
        end
    end
end
0.039520 seconds (629.95 k allocations: 9.811 MiB, 21.17% compilation time)

```

but when I use @threads as below, I have a longer execution time:

```julia
using .Threads

n=10000;
s=1
Y=zeros(n,n);

@time begin
    @threads for i in 1:n
        for m in 1:30
            if s == 1
                Y[i,i] += m;
            end
        end
    end
end
0.076712 seconds (1.80 M allocations: 30.648 MiB, 92.55% compilation time)

```

Any explanation or suggestion please?  
Thank you!

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 9, 2021, 1:25am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/2 "2021-07-09T01:25:36Z")

</div>

Threads can have have pretty high overhead – see for example

> <https://stackoverflow.com/questions/63933356/julia-threads-threads-slower-than-single-thread-performance/68310079#68310079>

> [@Julia Threads.@threads slower than single thread performance](https://discourse.julialang.org/t/julia-threads-threads-slower-than-single-thread-performance/46823):
>
> I am trying to solve numerically the heat equation in 1d: [[enter image description here]](https://i.stack.imgur.com/8GbwB.gif) I am using finite differences and I have some trouble using the @threads instruction in Julia. In particular below there are two version of the same code: the first one is single thread while the other uses @threads (they are identical apart from the @thread instruction) function heatSecLoop(;T::Float64) println("start") L = 1 ν = 0.5 Δt = 1e-6 Δx = 1e-3 Nt = ceil(Int, T/Δt ) …

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 9, 2021, 1:30am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/3 "2021-07-09T01:30:31Z")

</div>

Though actually one should switch to timing with `@btime` before drawing any serious conclusions from the above

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [July 9, 2021, 1:31am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/4 "2021-07-09T01:31:25Z")

</div>

While threads do have overhead, there are some other issues, you are using variables in a global scope, so their type can’t be inferred during compilation, you are also running in global scope, which can have issues, you are also including compilation time with your `@time` measurements, which shows that the threaded version spends 92% of the time compiling, compared to 21% for the non-threaded. A more fair comparison is this

```julia
function foo()
       n=10000
       s = 1
        for i in 1:n
                    for m in 1:30
                          if s == 1
                              Y[i,i] += m;
                          end
                      end
                  end
       end
foo (generic function with 1 method)

julia> foo()

julia> @time foo()
  0.033625 seconds (1.74 M allocations: 26.530 MiB)

julia> function foo2()
       n=10000
       s = 1
        @threads for i in 1:n
                      for m in 1:30
                          if s == 1
                              Y[i,i] += m;
                          end
                      end
                  end
                end
foo2 (generic function with 1 method)

julia> foo2()

julia> @time foo2()
  0.011409 seconds (1.74 M allocations: 26.532 MiB)

```

Sorry fo the crappy indenting

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [July 9, 2021, 1:40am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/6 "2021-07-09T01:40:42Z")

</div>

Here you go, should’ve included it on the first one

```julia
julia> @benchmark foo()
BechmarkTools.Trial: 282 samples with 1 evaluations.
 Range (min … max): 14.471 ms … 41.847 ms ┊ GC (min … max): 0.00% … 63.10%
 Time (median): 16.291 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 17.727 ms ± 5.913 ms ┊ GC (mean ± σ): 9.13% ± 15.23%

  ▂ ▇█▁
  █▆████▄▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▄▁▁▁▅▁▇▇▆ ▅
  14.5 ms Histogram: log(frequency) by time 40.4 ms <

 Memory estimate: 26.53 MiB, allocs estimate: 1738680.

julia> @benchmark foo2()
BechmarkTools.Trial: 761 samples with 1 evaluations.
 Range (min … max): 3.973 ms … 117.255 ms ┊ GC (min … max): 0.00% … 94.69%
 Time (median): 4.648 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 6.557 ms ± 14.078 ms ┊ GC (mean ± σ): 28.28% ± 12.44%

  █
  █▇▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▁▄▄ ▆
  3.97 ms Histogram: log(frequency) by time 113 ms <

 Memory estimate: 26.53 MiB, allocs estimate: 1738701.

```

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 9, 2021, 1:45am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/8 "2021-07-09T01:45:49Z")

</div>

Yeah, similarly:

```julia
julia> Threads.nthreads()
4

julia> @btime begin
          for i in 1:$n
              for m in 1:30
                  if $s == 1
                      $Y[i,i] += m
                  end
              end
          end
       end

  316.747 μs (0 allocations: 0 bytes)
julia> @btime begin
          Threads.@threads for i in 1:$n
              for m in 1:30
                  if $s == 1
                      $Y[i,i] += m
                  end
              end
          end
       end
  182.378 μs (20 allocations: 1.59 KiB)

```

where the `$` s are for interpolating global variables into the benchmark.  
You can obtain even better performance with LoopVectorization though:

```julia
julia> @btime begin
          @turbo for i in 1:$n
              for m in 1:30
                  Y[i,i] += m * ($s == 1)
              end
          end
       end
  163.157 μs (15 allocations: 448 bytes)

julia> @btime begin 
          @tturbo for i in 1:$n #Multi-threaded
              for m in 1:30
                  Y[i,i] += m * ($s == 1)
              end
          end
       end
  46.587 μs (15 allocations: 448 bytes)

```

`@turbo` is does single-threaded SIMD vectorization, `@tturbo` the same but with lightweight multithreading from Polyester.jl. Note that you have to eliminate the `if` statement (and effectively evaluate both sides of the branch) to use this, but it still comes out ahead.

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [July 9, 2021, 1:46am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/9 "2021-07-09T01:46:03Z")

</div>

I will always recommend taking a look at [Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/#kernel-functions) there you will find most of the pitfalls I explained and some others, leading to better perfomance and less issues while benchmarking.

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [July 9, 2021, 1:48am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/10 "2021-07-09T01:48:33Z")

</div>

Yeah, [https://github.com/JuliaSIMD/LoopVectorization.jl](https://github.com/JuliaSIMD/LoopVectorization.jl) , [https://github.com/JuliaSIMD/LoopVectorization.jl](https://github.com/JuliaSIMD/LoopVectorization.jl) and [GitHub - JuliaFolds/FLoops.jl: Fast sequential, threaded, and distributed for-loops for Julia—fold for humans™](https://github.com/JuliaFolds/FLoops.jl) are good resources for speeding things up.

---

<div class="post-metadata">

**Author:** ![Amro](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@Amro](https://discourse.julialang.org/u/Amro)\
**Post date:** [July 9, 2021, 2:01am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/11 "2021-07-09T02:01:30Z")

</div>

I am really confused when I am running my code

```julia
using .Threads

function foo()

n=10000;
s=1
Y=zeros(n,n);

    for i in 1:n
        for m in 1:30
            if s == 1
                Y[i,i] += m;
            end
        end
    end
end
julia> @time foo()
  0.162492 seconds (2 allocations: 762.940 MiB, 4.30% gc time)

julia> nthreads()
12
using .Threads

function foo2()

n=10000;
s=1
Y=zeros(n,n);

    @threads for i in 1:n
        for m in 1:30
            if s == 1
                Y[i,i] += m;
            end
        end
    end
end
julia> @time foo2()
  0.175901 seconds (63.74 k allocations: 766.888 MiB, 3.07% gc time, 10.05% compilation time)

```

 ![222](https://global.discourse-cdn.com/julialang/original/3X/4/2/42a5d499aa2d60f224309bf0198fa405fc1072e7.png)

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 9, 2021, 2:23am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/12 "2021-07-09T02:23:35Z")

</div>

The top one in your screenshot is counting compilation again. You should probably use `@btime` or `@benchmark` to avoid that happening by accident all the time.

---

<div class="post-metadata">

**Author:** ![Amro](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@Amro](https://discourse.julialang.org/u/Amro)\
**Post date:** [July 9, 2021, 2:39am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/16 "2021-07-09T02:39:56Z")

</div>

![44](https://global.discourse-cdn.com/julialang/original/3X/7/f/7fa2edb7f9a34a2d974a53ccca7847c3e4aeca69.png)  
Is this way in using @benchmark is correct?

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 9, 2021, 2:48am UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/17 "2021-07-09T02:48:23Z")

</div>

That seems to be allocating a massive amount of memory (700 MB!), so I suspect a global variable is sneaking in there somehow (perhaps if the function foo uses any variables that are not either explicitly passed to it or defined within it).

You might try copy-pasting the snippets I wrote above, which use $ for interpolation

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [July 9, 2021, 1:53pm UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/18 "2021-07-09T13:53:39Z")

</div>

The allocs are the nxn matrix, its quite large

```julia
a = zeros(10000,10000)
julia> varinfo()
  name size summary
  –––––––––––––––– ––––––––––– –––––––––––––––––––––––––––
  Base Module
  Core Module
  InteractiveUtils 256.530 KiB Module
  Main Module
  a 762.939 MiB 10000×10000 Matrix{Float64}
  ans 762.939 MiB 10000×10000 Matrix{Float64}

```

You can see that while it uses a lot of memory it uses very few allocations

---

<div class="post-metadata">

**Author:** ![Amro](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@Amro](https://discourse.julialang.org/u/Amro)\
**Post date:** [July 9, 2021, 2:37pm UTC](https://discourse.julialang.org/t/question-for-lower-performance-by-using-threads-in-for-loop/64315/19 "2021-07-09T14:37:17Z")

</div>

Do you think my way in using @benchmark is correct because the results of both foo and foo2 functions are close to each other, however, foo2 should gives better performance?
