# Intriguing performance comparison results.... threads, comprehensions, etc

**URL:** <https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093>\
**Category:** Performance\
**Created:** [January 23, 2025, 12:37am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093 "2025-01-23T00:37:51Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 12:37am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/1 "2025-01-23T00:37:51Z")

</div>

Can anyone provide me with intuition for the poor performance of the dynamic threads for loop in the example below? (all second run times) Thanks!

```julia
using LinearAlgebra, .Threads

function test( Y, ::Val{1} )
    X = similar( Y )
    for n ∈ eachindex( X )
        X[n] = inv( Y[n] )
    end
    return X
end

function test( Y, ::Val{2} )
    return inv.( Y )
end

function test( Y, ::Val{3} )
    X = similar( Y )
    @threads :static for n ∈ eachindex( X )
        X[n] = inv( Y[n] )
    end
    return X
end

function test( Y, ::Val{4} )
    X = similar( Y )
    @threads :dynamic for n ∈ eachindex( X )
        X[n] = inv( Y[n] )
    end
    return X
end

function test( Y, ::Val{5} )
    return map( inv, Y )
end

function test( Y, :: Val{6} )
    return [inv(y) for y ∈ Y]
end

generate( N, n ) = [randn( n, n ) for i ∈ 1:N]

function runme( N, n )
    X = generate( N, n )
    for i ∈ 1:6
        @time test( X, Val( i ) )
    end
    for i ∈ 1:6
        print( "$i: " )
        @time test( X, Val( i ) )
    end

end

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/a/eacd9e9e33cce98ac26fc10a0e54758f38181cc0.png)

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [January 23, 2025, 2:19am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/2 "2025-01-23T02:19:21Z")

</div>

> [@Joris\_Pinkse](#):
>
> the poor performance

How do you conclude that your results show a poor performance? What is your reference/ expected result?

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [January 23, 2025, 2:52am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/3 "2025-01-23T02:52:43Z")

</div>

What was `Threads.nthreads()`?

---

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 3:00am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/4 "2025-01-23T03:00:23Z")

</div>

Method 4 takes more than twice as long as the other five methods.

---

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 3:02am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/5 "2025-01-23T03:02:51Z")

</div>

1. (#hyperthreads, difference smaller when using #physical cores, which I tend to use in practice}

---

<div class="post-metadata">

**Author:** ![torrance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torrance/32/38990_2.png) [@torrance](https://discourse.julialang.org/u/torrance)\
**Post date:** [January 23, 2025, 4:30am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/6 "2025-01-23T04:30:59Z")

</div>

Some things to note:

- `inv()` calls out to BLAS and is already threaded. So the threaded options here will only add extra overhead.
- Dynamic scheduling will have extra overhead. In cases where `n` is small and `N’ is especially large, this overhead will be proportionately larger.

In my own tests with (N=100, n=1000), (N=1000, n=1000), (N=1000000, n=100) I see basically all tests giving the same approximate runtime. What N,n were you using?

---

<div class="post-metadata">

**Author:** ![karei](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/karei/32/214809_2.png) [@karei](https://discourse.julialang.org/u/karei)\
**Post date:** [January 23, 2025, 6:30am UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/7 "2025-01-23T06:30:18Z")

</div>

As torrance said, the BLAS function itself is threaded. If you use multiple Julia threads, the total threads have a complicated relationship with BLAS threads and Julia threads. If the total number of threads exceeds your CPU cores, you will experience poor performance.

This document clearly explains such an issue:

> **[Pinning BLAS Threads · ThreadPinning.jl](https://carstenbauer.github.io/ThreadPinning.jl/stable/examples/ex_blas/#Beware:-Interaction-between-Julia-threads-and-BLAS-threads)**
>
> Documentation for ThreadPinning.jl.

---

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 3:02pm UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/8 "2025-01-23T15:02:17Z")

</div>

This one is for 1000, 1000.

---

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 3:05pm UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/9 "2025-01-23T15:05:22Z")

</div>

Thanks @karei . Indeed. I’m just surprised by the chasm between dynamic threads and the other methods. I’m familiar with the ThreadPinning package, though I haven’t found that in my application there was a significant performance increase.

The reason I’m looking into this is performance degradation in my package from one Julia version to the next and the fact that now reducing the number of cores often improves runtime.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [January 23, 2025, 5:48pm UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/10 "2025-01-23T17:48:00Z")

</div>

> [@Joris\_Pinkse](#):
>
> the fact that now reducing the number of cores often improves runtime.

Less cores → less cache pressure → better performance  
Less threads → less thread switching overhead → better performance

A high number of threads is only useful if:

- your code does not allocate much, low GC overhead
- the execution time within each thread is high
- the memory needed for each thread is low
- your CPU has lots of cache and/or a very high memory bandwidth

---

<div class="post-metadata">

**Author:** ![Joris\_Pinkse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joris_pinkse/32/216398_2.png) [@Joris\_Pinkse](https://discourse.julialang.org/u/Joris_Pinkse)\
**Post date:** [January 23, 2025, 6:58pm UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/11 "2025-01-23T18:58:01Z")

</div>

Thanks @ufechner7 . I understand all of that, but am still surprised by the 2.5 times performance cliff that only shows with _dynamic_ threads. I had expected there to be a cost but didn’t realize creating / using extra threads was _that_ costly.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [January 23, 2025, 7:03pm UTC](https://discourse.julialang.org/t/intriguing-performance-comparison-results-threads-comprehensions-etc/125093/12 "2025-01-23T19:03:56Z")

</div>

If you need cheap threads, use [GitHub - JuliaSIMD/Polyester.jl: The cheapest threads you can find!](https://github.com/JuliaSIMD/Polyester.jl)

An additional aspect is the fact that the CPU clock goes down the more CPU cores you use… This effect is very different depending on the CPU, though.
