# Ideal number of BLAS threads

**URL:** <https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197>\
**Category:** General Usage\
**Tags:** blas, multithreading, linearalgebra\
**Created:** [April 7, 2022, 10:30pm UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197 "2022-04-07T22:30:16Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![fipelle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fipelle/32/4772_2.png) [@fipelle](https://discourse.julialang.org/u/fipelle)\
**Post date:** [April 7, 2022, 10:30pm UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/1 "2022-04-07T22:30:16Z")

</div>

Hi,

I am planning to run a heavy series of matrix operations on a cluster and I was wondering what’s the ideal number of BLAS threads I should use. Assume that I have a machine with 20 cores and 8-16 GB of RAM per core. Does it make sense to use more BLAS threads than cores? If so, is there any rule of thumb I should follow?

PS: I am migrating towards Octavian.jl. How would it change if I were to use it instead of LinearAlgebra.jl?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 8, 2022, 3:40am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/2 "2022-04-08T03:40:43Z")

</div>

> [@fipelle](#):
>
> Does it make sense to use more BLAS threads than cores?

No.  
Octavian (and MKL) should also both use at most 1 thread per core.  
Octavian runs on Julia’s threads, rather than being controlled by `BLAS.set_num_threads()`.

---

<div class="post-metadata">

**Author:** ![photor](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/photor/32/14343_2.png) [@photor](https://discourse.julialang.org/u/photor)\
**Post date:** [April 8, 2022, 3:49am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/3 "2022-04-08T03:49:04Z")

</div>

Just use

```julia
julia -t (number of physical cores)

```

to start julia for Octavian.jl

---

<div class="post-metadata">

**Author:** ![fipelle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fipelle/32/4772_2.png) [@fipelle](https://discourse.julialang.org/u/fipelle)\
**Post date:** [April 15, 2022, 4:37am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/4 "2022-04-15T04:37:19Z")

</div>

I have a follow-up question. I have noticed that on a cluster with 1vCPU per core and a total of 64 cores `BLAS.set_num_threads()` sets a maximum of 32 cores (not 64). Would you please clarify what is the reasoning behind it? I have noticed with `lscpu` that there are 32 cores per socket - is that it? Also, is there any advantage in using a different ratio of vCPU per core?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 26, 2022, 1:03am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/5 "2022-04-26T01:03:26Z")

</div>

Yes, that is probably it.  
No longer using Hwloc.jl means CPUSummary.jl doesn’t know the number of machines.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [April 26, 2022, 5:01am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/6 "2022-04-26T05:01:32Z")

</div>

He says that `BLAS.set_num_threads` is limited to max 32 threads which shouldn’t be related to Hwloc.jl or CPUSummary.jl at all. Instead it’s just that Julia \<= 1.7 has a built in hard limit of max 32 OpenBLAS threads. To work around it you can either use a different BLAS or switch to Julia \> 1.7. (or change the limit manually and compile Julia yourself).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 26, 2022, 5:48am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/7 "2022-04-26T05:48:06Z")

</div>

> [@carstenbauer](#):
>
> He says that `BLAS.set_num_threads` is limited to max 32 threads which shouldn’t be related to Hwloc.jl or CPUSummary.jl at all.

Ha, sorry – I should have read more closely / was still thinking about the earlier discussion mentioning Octavian.  
It’s threads would be limited by this.

Yes, for `BLAS.set_num_threads`/base Julia, use Julia \> 1.7.  
If on x86, you can also consider trying MKL.jl. MKL should perform well on Intel, and recent versions may at least be competitive with OpenBLAS on AMD.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [April 26, 2022, 5:56am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/8 "2022-04-26T05:56:58Z")

</div>

If you’re curious, you can also be (most likely) the first person besides me to try [GitHub - carstenbauer/BLISBLAS.jl: BLIS-pendant of MKL.jl](https://github.com/carstenbauer/BLISBLAS.jl). [BLIS](https://github.com/flame/blis) should run well and can outperform OpenBLAS, in particular on AMD, in some cases.

Note though that BLIS only provides BLAS and no LAPACK (OpenBLAS will still be used for this).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 26, 2022, 4:00pm UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/9 "2022-04-26T16:00:47Z")

</div>

Neat.  
I’ve tried BLIS\_jll before, but it performed extremely poorly at small sizes, especially when multithreading. At the time at least, they didn’t have an equivalent of OpenBLAS’s multithreaded gemm threshold (so they were always threading, regardless of how small).

> **[GitHub - JuliaLinearAlgebra/BLASBenchmarksCPU.jl: Benchmark BLAS Libraries](https://github.com/JuliaLinearAlgebra/BLASBenchmarksCPU.jl)**
>
> Benchmark BLAS Libraries. Contribute to JuliaLinearAlgebra/BLASBenchmarksCPU.jl development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [April 27, 2022, 12:49am UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/10 "2022-04-27T00:49:44Z")

</div>

> [@fipelle](#):
>
> I have noticed that on a cluster with 1vCPU per core and a total of 64 cores `BLAS.set_num_threads()` sets a maximum of 32 cores (not 64). Would you please clarify what is the reasoning behind it?

Apart to above suggestions you may also take a look at Julia 1.8 [[BLAS performance testing for Julia 1.8](https://discourse.julialang.org/t/blas-performance-testing-for-julia-1-8/69520)]. I have not followed it recently, however as I understand it currently, there might be some significant changes and automation associated with BLAS and 1.8 release.

@carstenbauer Would BLISBLAS.jl work on Neoverse N1 (Ampere Altra)? I have heard some words that on this particular CPU BLIS might be one of the most favorable options.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [April 27, 2022, 2:43pm UTC](https://discourse.julialang.org/t/ideal-number-of-blas-threads/79197/11 "2022-04-27T14:43:22Z")

</div>

I am confirming it seems to work without any problems on Neoverse-N1. As for the performance, I was not able to do any in depth testing. Some preliminary info based on the code from the repository pls find below.  
I am recalling I had a discussion about BLIS on ARM (Neoverse-N1) a few months ago and it looked challenging for me to make it work then. Thank you for putting all this together.

> lscpu  
> Architecture: aarch64  
> CPU op-mode(s): 32-bit, 64-bit  
> Byte Order: Little Endian  
> CPU(s): 4  
> On-line CPU(s) list: 0-3  
> Thread(s) per core: 1  
> Core(s) per socket: 4  
> Socket(s): 1  
> NUMA node(s): 1  
> Vendor ID: ARM  
> Model: 1  
> Model name: Neoverse-N1  
> Stepping: r3p1  
> BogoMIPS: 50.00  
> NUMA node0 CPU(s): 0-3  
> Vulnerability Itlb multihit: Not affected  
> Vulnerability L1tf: Not affected  
> Vulnerability Mds: Not affected  
> Vulnerability Meltdown: Not affected  
> Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl  
> Vulnerability Spectre v1: Mitigation; \_\_user pointer sanitization  
> Vulnerability Spectre v2: Mitigation; CSV2, BHB  
> Vulnerability Srbds: Not affected  
> Vulnerability Tsx async abort: Not affected  
> Flags: fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm lrcpc dcpop asimddp ssbs

> ./julia -t auto  
> julia\> versioninfo()  
> Julia Version 1.7.0  
> Commit 3bf9d17731 (2021-11-30 12:12 UTC)  
> Platform Info:  
> OS: Linux (aarch64-unknown-linux-gnu)  
> CPU: unknown  
> WORD\_SIZE: 64  
> LIBM: libopenlibm  
> LLVM: libLLVM-12.0.1 (ORCJIT, neoverse-n1)

BLIS:

> julia\> using LinearAlgebra  
> julia\> BLAS.get\_config()  
> LinearAlgebra.BLAS.LBTConfig  
> Libraries:  
> └ [ILP64] libopenblas64\_.so  
> julia\> using BLISBLAS  
> julia\> BLAS.get\_config()  
> LinearAlgebra.BLAS.LBTConfig  
> Libraries:  
> ├ [ILP64] libopenblas64\_.so  
> └ [ILP64] libblis.so  
> julia\> using BenchmarkTools  
> julia\> A = rand(1000,1000); B = rand(1000,1000);  
> julia\> @btime $A \* $B;  
> 98.963 ms (2 allocations: 7.63 MiB)

OPENBLAS:

> julia\> using LinearAlgebra  
> julia\> BLAS.get\_config()  
> LinearAlgebra.BLAS.LBTConfig  
> Libraries:  
> └ [ILP64] libopenblas64\_.so  
> julia\> using BenchmarkTools  
> julia\> A = rand(1000,1000); B = rand(1000,1000);  
> julia\> @btime $A \* $B;  
> 29.436 ms (2 allocations: 7.63 MiB)
