# Multithreaded code on beefy computer runs just as fast as serial code on M1 Mac

**URL:** <https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210>\
**Category:** Performance\
**Tags:** performance, parallel, differentialequation\
**Created:** [January 26, 2022, 2:30am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210 "2022-01-26T02:30:17Z")\
**Posts on this page:** 1\
**Showing post:** 32

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [January 26, 2022, 8:06pm UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/32 "2022-01-26T20:06:47Z")

</div>

> [@ash](#):
>
> MKL is 4 times slower in parallel!!

Disclaimer: I didn’t read the entire thread / didn’t follow too closely.

Did you benchmark everything with `BLAS.set_num_threads(1)`? If not, note that you need to take when using MKL and, at the same time, multiple Julia threads (same can be said about OpenBLAS). I find that `MKL_NUM_THREADS` defaults to the number of cores of the system and since this number is used _per Julia thread_ you will readily have too many MKL threads running (and thus oversubscribe your cores). See [Matrix multiplication is slower when multithreading in Julia - #12 by carstenbauer](https://discourse.julialang.org/t/matrix-multiplication-is-slower-when-multithreading-in-julia/56227/12) for more.

TLDR: I suggest you benchmark everything with `BLAS.set_num_threads(1)` (independent of whether you use MKL or OpenBLAS).

---

_[View the full topic](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210)._
