# OpenBLAS vs MKL

**URL:** <https://discourse.julialang.org/t/openblas-vs-mkl/17362>\
**Category:** General Usage\
**Tags:** mkl\
**Created:** [November 9, 2018, 8:19pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362 "2018-11-09T20:19:19Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [November 9, 2018, 8:19pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/1 "2018-11-09T20:19:19Z")

</div>

I just build Julia 1.0.1 on a cluster and linked it against MKL (2019). As a simple benchmark, I compared the performance of squaring a `1000 x 1000` matrix against the Julia binaries from the website (OpenBLAS). I remember I did this for 0.6.4 at some point and found MKL to be faster by 30% or so. However, I’m blown away by the difference I found this time:

MKL:

```julia
julia> versioninfo()
Julia Version 1.0.1
Commit 0d713926f8* (2018-09-29 19:05 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Xeon(R) Gold 6148 CPU @ 2.40GHz
  WORD_SIZE: 64
  LIBM: libimf
  LLVM: libLLVM-6.0.0 (ORCJIT, skylake)

julia> using LinearAlgebra; LinearAlgebra.versioninfo()
BLAS: libmkl_rt
LAPACK: libmkl_rt

julia> using BenchmarkTools

julia> A = rand(1000,1000);

julia> @btime $A*$A;
  1.926 ms (2 allocations: 7.63 MiB)

```

OpenBLAS:

```julia
julia> versioninfo()
Julia Version 1.0.1
Commit 0d713926f8 (2018-09-29 19:05 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Xeon(R) Gold 6148 CPU @ 2.40GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.0 (ORCJIT, skylake)

julia> using LinearAlgebra; LinearAlgebra.versioninfo()
BLAS: libopenblas (USE64BITINT DYNAMIC_ARCH NO_AFFINITY SkylakeX MAX_THREADS=16)
LAPACK: libopenblas64_

julia> using BenchmarkTools

julia> A = rand(1000,1000);

julia> @btime $A*$A;
  7.905 ms (2 allocations: 7.63 MiB)

julia> @btime $A*$A;
  7.124 ms (2 allocations: 7.63 MiB)

```

---

<div class="post-metadata">

**Author:** ![davidbp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidbp/32/463_2.png) [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Post date:** [November 9, 2018, 8:54pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/2 "2018-11-09T20:54:44Z")

</div>

It would be nice for you to share this but for matrices of different sizes. `(10,10), (100,100), (500,500), (1000,1000), (10000,10000)` maybe MKL is very good for some sizes but not others. In any case a 4x is relevant but not an insane number. Unless your code is completely dominated by matrix multiplications it won’t be such a big deal.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [November 9, 2018, 9:39pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/3 "2018-11-09T21:39:17Z")

</div>

The cluster has avx-512:

```julia
SkylakeX

```

Also, see [Intel Ark Xeon(R) Gold 6148](https://ark.intel.com/products/120489/Intel-Xeon-Gold-6148-Processor-27-5M-Cache-2-40-GHz-)  
The latest OpenBLAS has finally gotten some support for avx512 dgemm, but their kernels are still far from optimal. Julia 1.0.1 does not come with the latest OpenBLAS.

I see a similar difference on my Skylake-X cpu. OpenBLAS does far better on avx2 architectures, like Haswell.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [November 10, 2018, 1:37am UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/4 "2018-11-10T01:37:14Z")

</div>

> [@davidbp](#):
>
> It would be nice for you to share this but for matrices of different sizes. `(10,10), (100,100), (500,500), (1000,1000), (10000,10000)` maybe MKL is very good for some sizes but not others.

![openblas_vs_mkl](https://global.discourse-cdn.com/julialang/original/3X/5/f/5f5dbed941a2a57409ff9a001e66fe08ee84f8e0.png) ![timings](https://global.discourse-cdn.com/julialang/original/3X/4/a/4a6e6ae7c7f26634289dc3e59f2e13a778d3b9cc.png)

> [@davidbp](#):
>
> In any case a 4x is relevant but not an insane number. Unless your code is completely dominated by matrix multiplications it won’t be such a big deal.

My code is completely dominated by matrix multiplications.

---

<div class="post-metadata">

**Author:** ![e3c6](https://avatars.discourse-cdn.com/v4/letter/e/e79b87/32.png) [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Post date:** [January 15, 2020, 7:56pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/5 "2020-01-15T19:56:57Z")

</div>

Wow that’s a big difference. After one year the picture remains the same?

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [January 15, 2020, 8:08pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/6 "2020-01-15T20:08:54Z")

</div>

Most of the performance gap is likely due to [BLAS threads should default to physical not logical core count? · Issue #33409 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/33409)

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [January 15, 2020, 11:09pm UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/7 "2020-01-15T23:09:38Z")

</div>

MKL might also switch over to Strassen’s method or similar sub n^3 methods when the size gets big.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 16, 2020, 1:01am UTC](https://discourse.julialang.org/t/openblas-vs-mkl/17362/8 "2020-01-16T01:01:30Z")

</div>

I don’t think it does based on `LinearAlgebra.peakflops(N)`.
