# MKL slower than openblas in intel cpu

**URL:** <https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169>\
**Category:** Performance\
**Tags:** mkl, linearalgebra\
**Created:** [February 28, 2022, 3:00am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169 "2022-02-28T03:00:50Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jae-Mo\_Lihm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jae-mo_lihm/32/20145_2.png) [@Jae-Mo\_Lihm](https://discourse.julialang.org/u/Jae-Mo_Lihm)\
**Post date:** [February 28, 2022, 3:00am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/1 "2022-02-28T03:00:50Z")

</div>

I found MKL is ~1.5 times slower than openblas for matrix-matrix multiplication (both without multithreading). Since I’m using intel CPU, I expected MKL to be faster.

Is there anything obviously wrong in this benchmark? Or is it not unexpected for MKL to be slower than openblas?

```julia
using LinearAlgebra, StaticArrays, BenchmarkTools
N = 400
M = 5000
A = rand(ComplexF64, N, M);
x = rand(ComplexF64, M, 3);
y = rand(ComplexF64, N, 3);
LinearAlgebra.BLAS.set_num_threads(1)
@btime mul!($y, $A, $x);
# 3.323 ms (0 allocations: 0 bytes)

using MKL
LinearAlgebra.BLAS.set_num_threads(1)
@btime mul!($y, $A, $x);
# 5.168 ms (0 allocations: 0 bytes)

```

```julia
julia> versioninfo()
Julia Version 1.7.0
Commit 3bf9d17731 (2021-11-30 12:12 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Xeon(R) Gold 6230 CPU @ 2.10GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-12.0.1 (ORCJIT, cascadelake)
Environment:
  JULIA = /home/jmlim/appl/julia-1.7.0/bin/julia

```

---

<div class="post-metadata">

**Author:** ![Jae-Mo\_Lihm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jae-mo_lihm/32/20145_2.png) [@Jae-Mo\_Lihm](https://discourse.julialang.org/u/Jae-Mo_Lihm)\
**Post date:** [February 28, 2022, 3:03am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/2 "2022-02-28T03:03:48Z")

</div>

I found the performance is specific to that matrix size. (very narrow columns)

```julia
using LinearAlgebra, StaticArrays, BenchmarkTools
N = 300
M = 300
K = 300
A = rand(ComplexF64, N, M);
x = rand(ComplexF64, M, K);
y = rand(ComplexF64, N, K);
LinearAlgebra.BLAS.set_num_threads(1)
@btime mul!($y, $A, $x);
# 2.209 ms (0 allocations: 0 bytes)

using MKL
LinearAlgebra.BLAS.set_num_threads(1)
@btime mul!($y, $A, $x);
# 2.015 ms (0 allocations: 0 bytes)

```

---

<div class="post-metadata">

**Author:** ![photor](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/photor/32/14343_2.png) [@photor](https://discourse.julialang.org/u/photor)\
**Post date:** [February 28, 2022, 3:42am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/3 "2022-02-28T03:42:25Z")

</div>

So, it’s normal.

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [February 28, 2022, 5:31am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/4 "2022-02-28T05:31:45Z")

</div>

- With MKL.jl v0.5.0? (released this 2 days ago)
- Could you please test with the new [Julia v1.8.0-beta1](https://discourse.julialang.org/t/julia-v1-8-0-beta1-is-now-available/77166)

---

<div class="post-metadata">

**Author:** ![Seif\_Shebl](https://avatars.discourse-cdn.com/v4/letter/s/eada6e/32.png) [@Seif\_Shebl](https://discourse.julialang.org/u/Seif_Shebl)\
**Post date:** [February 28, 2022, 7:02pm UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/5 "2022-02-28T19:02:23Z")

</div>

I can’t reproduce.

```julia
  7.244 ms (0 allocations: 0 bytes) OpenBLAS
  4.748 ms (0 allocations: 0 bytes) MKL (0.5.0)

```

On an 11 days old Julia 1.8.0-DEV with Intel CPU:

```julia
Julia Version 1.8.0-DEV.1572
Commit 7889b2a6a2 (2022-02-16 21:17 UTC)
Platform Info:
  OS: Windows (x86_64-w64-mingw32)
  CPU: 8 × Intel(R) Core(TM) i7-4790K CPU @ 4.00GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-13.0.1 (ORCJIT, haswell)
  Threads: 8 on 8 virtual cores

```

---

<div class="post-metadata">

**Author:** ![Jae-Mo\_Lihm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jae-mo_lihm/32/20145_2.png) [@Jae-Mo\_Lihm](https://discourse.julialang.org/u/Jae-Mo_Lihm)\
**Post date:** [March 2, 2022, 1:22am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/6 "2022-03-02T01:22:04Z")

</div>

Yes, I get basically the same timings with v1.8.0-beta1 + MKL 0.5.0.

---

<div class="post-metadata">

**Author:** ![Jae-Mo\_Lihm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jae-mo_lihm/32/20145_2.png) [@Jae-Mo\_Lihm](https://discourse.julialang.org/u/Jae-Mo_Lihm)\
**Post date:** [March 2, 2022, 1:23am UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/7 "2022-03-02T01:23:05Z")

</div>

Thanks for running the benchmark. The problem seems to be specific to my CPU.

To be clear, did you run the one in the OP ( (400, 5000) \* (5000, 3)) or the one in the reply ( (300, 300) \* (300, 300) )?

---

<div class="post-metadata">

**Author:** ![Seif\_Shebl](https://avatars.discourse-cdn.com/v4/letter/s/eada6e/32.png) [@Seif\_Shebl](https://discourse.julialang.org/u/Seif_Shebl)\
**Post date:** [March 2, 2022, 3:48pm UTC](https://discourse.julialang.org/t/mkl-slower-than-openblas-in-intel-cpu/77169/8 "2022-03-02T15:48:01Z")

</div>

The results above are for the sizes in the OP, (400, 5000) \times(5000, 3), and the following are the results for (300, 300)\times(300, 300).

```julia
LinearAlgebra.BLAS.set_num_threads(1)
...
  3.819 ms (0 allocations: 0 bytes) OpenBLAS
  3.617 ms (0 allocations: 0 bytes) MKL (0.5.0)

```
