# BLAS vs Threads on a cluster

**URL:** <https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332>\
**Category:** Performance\
**Created:** [April 22, 2024, 11:47am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332 "2024-04-22T11:47:15Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![SpuriousEigenstate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/spuriouseigenstate/32/208390_2.png) [@SpuriousEigenstate](https://discourse.julialang.org/u/SpuriousEigenstate)\
**Post date:** [April 22, 2024, 11:47am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/1 "2024-04-22T11:47:15Z")

</div>

I’m developing a script which, at its core, is mainly diagonalizing a large number of Hermitian matrices, i.e., `eigen(Hermitian(A))`. I want to run the code on a cluster (on a single machine, 96 threads) and am looking into multithreading it using `Threads.@threads`.

In testing on my desktop (16 threads), I noticed that the `LinearAlgebra` functions automatically make use of multiple threads. So my question is: is it better (faster) to let all the threads be used by `eigen`, or to restrict `eigen` to a single thread and parallelize using `Threads`? Or somewhere in between?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 22, 2024, 12:09pm UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/2 "2024-04-22T12:09:08Z")

</div>

Generally you want to set BLAS to a single thread if you can paralellize at a higher level easily. BLAS multithreading doesn’t give you perfect scaling.

---

<div class="post-metadata">

**Author:** ![SpuriousEigenstate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/spuriouseigenstate/32/208390_2.png) [@SpuriousEigenstate](https://discourse.julialang.org/u/SpuriousEigenstate)\
**Post date:** [April 22, 2024, 12:29pm UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/3 "2024-04-22T12:29:50Z")

</div>

Thanks! A simple test shows that you are right:

```julia
using LinearAlgebra

A = rand(800,800)

@time Threads.@threads for i in 1:30
    eigen(A)
end

BLAS.set_num_threads(1)
@time Threads.@threads for i in 1:30
    eigen(A)
end

```

output:

```julia
 12.973553 seconds (31.24 k allocations: 613.706 MiB, 0.16% gc time, 3.23% compilation time)
  5.900912 seconds (22.85 k allocations: 613.151 MiB, 0.23% gc time, 9.74% compilation time)

```

---

<div class="post-metadata">

**Author:** ![abraemer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abraemer/32/51403_2.png) [@abraemer](https://discourse.julialang.org/u/abraemer)\
**Post date:** [April 23, 2024, 5:18am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/4 "2024-04-23T05:18:19Z")

</div>

Also note that there is an interaction between BLAS threads and julia threads that depends on whether you use MKL or the default OpenBLAS. In short:

- with MKL: total threads used = julia threads x BLAS threads
- with OpenBLAS: total threads used = julia threads + BLAS threads (BLAS will compute “jobs” sequentially but use threads within each “job”)

More details explanation can be found here:

> **[Julia Threads + BLAS Threads · ThreadPinning.jl](https://carstenbauer.github.io/ThreadPinning.jl/dev/explanations/blas/)**
>
> Documentation for ThreadPinning.jl.

(Also using ThreadPinning.jl might improve performance if you use a cluster)

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 23, 2024, 6:42am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/5 "2024-04-23T06:42:17Z")

</div>

When I wrote the docs section about this I buried it at the end of the performance tips, but maybe there is a better place?

[https://docs.julialang.org/en/v1/manual/performance-tips/#man-multithreading-linear-algebra](https://docs.julialang.org/en/v1/manual/performance-tips/#man-multithreading-linear-algebra)

---

<div class="post-metadata">

**Author:** ![abraemer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abraemer/32/51403_2.png) [@abraemer](https://discourse.julialang.org/u/abraemer)\
**Post date:** [April 23, 2024, 7:20am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/6 "2024-04-23T07:20:27Z")

</div>

I think the Performance tips section is a reasonable place in principle. I usually link to the explanation in ThreadPinning.jl because it is a bit more detailed and also highlights the different behavior of MKL.jl (which I deem a very surprising footgun…).

Maybe a lesson here could be that the Performance Tips section got a bit too long to be unstructured? Maybe we could make some sections like “Type related stuff”, “Function related stuff”, “Numerics”, “Miscellaneous”. Maybe lead with a section “Most common and severe performance pitfalls”.

For the specific performance tip about OpenBLAS, I think a bit more emphasis on the (to me at surprising) behavior of `OPENBLAS_NUM_THREADS=N>1` would be good. The current “There is just one OpenBLAS thread pool shared among all Julia threads.” does not scream to me “your operations will essentially be done serially (but a bit faster)”.

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 23, 2024, 8:06am UTC](https://discourse.julialang.org/t/blas-vs-threads-on-a-cluster/113332/7 "2024-04-23T08:06:29Z")

</div>

> [@abraemer](#):
>
> Maybe we could make some sections like “Type related stuff”, “Function related stuff”, “Numerics”, “Miscellaneous”.

I’ve been wanting to do that for a while, but the Documenter format doesn’t make such subsections visible, so I’m not sure how to proceed

> [@abraemer](#):
>
> For the specific performance tip about OpenBLAS, I think a bit more emphasis on the (to me at surprising) behavior of `OPENBLAS_NUM_THREADS=N>1` would be good. The current “There is just one OpenBLAS thread pool shared among all Julia threads.” does not scream to me “your operations will essentially be done serially (but a bit faster)”.

Feel free to submit a PR, I didn’t write this part because I was an expert, only because I was frustrated not to find it anywhere official ^^
