# @threads for loop performance

**URL:** <https://discourse.julialang.org/t/threads-for-loop-performance/51641>\
**Category:** Performance\
**Created:** [December 11, 2020, 12:05am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641 "2020-12-11T00:05:46Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ankit13](https://avatars.discourse-cdn.com/v4/letter/a/67e7ee/32.png) [@ankit13](https://discourse.julialang.org/u/ankit13)\
**Post date:** [December 11, 2020, 12:05am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/1 "2020-12-11T00:05:46Z")

</div>

Hi all,

Why my code is running slower when I use more number of threads in @threads for loop? ![run](https://global.discourse-cdn.com/julialang/original/3X/a/2/a2fb13c3794af32f64842eb7c80a8a3ba664df56.png)

I would like to know, if this is a standard behavior for threaded parallelization?

This is a sample code that will reproduce the behavior.

```julia
function check(n )
    y = [one( Array{ComplexF64}(undef,2,2) ) for i=1:n]
    t = [zero( Array{ComplexF64}(undef,2,2) ) for i=1:n]

   Threads.@threads for i=1:n
        for j=1:n
            for k=1:n
                t[i] += y[k]
            end
        end
    end
end

@time check(400)

```

Thanks.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [December 11, 2020, 12:14am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/2 "2020-12-11T00:14:45Z")

</div>

Can you post a minimal working example of the code? Otherwise, it’s hard to help.

---

<div class="post-metadata">

**Author:** ![ankit13](https://avatars.discourse-cdn.com/v4/letter/a/67e7ee/32.png) [@ankit13](https://discourse.julialang.org/u/ankit13)\
**Post date:** [December 11, 2020, 12:32am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/3 "2020-12-11T00:32:33Z")

</div>

I did not post it because it is quite big. However, I can explain it here. Overall what I am doing is following: I have an 1D array (length 10000) of 2x2 complex matrices and I apply some functions on theses matrices using the threaded for loop where each thread takes one element of the 1D array.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [December 11, 2020, 12:52am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/4 "2020-12-11T00:52:20Z")

</div>

How many physical cores does your computer have?

I’d also use `SMatrix` from StaticArrays.jl for 2x2 complex matrices.

---

<div class="post-metadata">

**Author:** ![ankit13](https://avatars.discourse-cdn.com/v4/letter/a/67e7ee/32.png) [@ankit13](https://discourse.julialang.org/u/ankit13)\
**Post date:** [December 11, 2020, 1:01am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/5 "2020-12-11T01:01:32Z")

</div>

I have 28 cores. I will try static arrays.

---

<div class="post-metadata">

**Author:** ![Jan\_Kybic1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jan_kybic1/32/20184_2.png) [@Jan\_Kybic1](https://discourse.julialang.org/u/Jan_Kybic1)\
**Post date:** [December 11, 2020, 8:57am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/6 "2020-12-11T08:57:30Z")

</div>

In your case, the problem seems to be related to your use of an Array of 2D matrices, which leads to a lot of allocations. Using `.+=` reduces allocations by doing the update in place. Using 3D matrices as follows

```julia
    y = ones(ComplexF64,(2,2,n))
    t = zeros(ComplexF64,(2,2,n))

```

is 10x faster on my computer and I get an additional speed-up from multithreading.

Interestingly, I also found recently that the performance deteriorated when I added `Threads.@threads` to my main `for` loop. In my case, the solution was to switch off multithreading in BLAS by calling `BLAS.set_num_threads(1)`. From that point on, I started to see a speed-up.  
It would be useful if we could at least get a warning in such cases.

Yours,

Jan

---

<div class="post-metadata">

**Author:** ![ankit13](https://avatars.discourse-cdn.com/v4/letter/a/67e7ee/32.png) [@ankit13](https://discourse.julialang.org/u/ankit13)\
**Post date:** [December 11, 2020, 11:57am UTC](https://discourse.julialang.org/t/threads-for-loop-performance/51641/7 "2020-12-11T11:57:54Z")

</div>

Okay, I will try. Thanks for the help.
