# Benchmark MATLAB & Julia for Matrix Operations

**URL:** <https://discourse.julialang.org/t/benchmark-matlab-julia-for-matrix-operations/2000>\
**Category:** Performance\
**Created:** [February 9, 2017, 12:50pm UTC](https://discourse.julialang.org/t/benchmark-matlab-julia-for-matrix-operations/2000 "2017-02-09T12:50:37Z")\
**Posts on this page:** 1\
**Showing post:** 92

<div class="post-metadata">

**Author:** ![barche](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barche/32/79_2.png) [@barche](https://discourse.julialang.org/u/barche)\
**Post date:** [February 14, 2017, 1:49pm UTC](https://discourse.julialang.org/t/benchmark-matlab-julia-for-matrix-operations/2000/92 "2017-02-14T13:49:37Z")

</div>

To add another datapoint, here are the results on a 32-core node on our cluster, with and without threading and comparing OpenBLAS and MKL:  
[https://github.com/barche/julia-blas-benchmarks/blob/master/BenchmarkResults.ipynb](https://github.com/barche/julia-blas-benchmarks/blob/master/BenchmarkResults.ipynb)

I also reran the HPL linpack test, here are the results:

- Standard HPL OpenBLAS, 32 MPI processes on a single node: **757 Gflops**
- Standard HPL MKL, 32 MPI processes on a single node: **788 Gflops**
- Intel HPL MKL, 32 MPI processes on a single node: **814 Gflops**
- Intel HPL MKL, 2 MPI processes with 16 threads each on a single node: **963 Gflops**

From both tests it seems clear to me that MKL wins when threading enters into the equation, but single-core performance is much closer, with the possible exception of the Cholesky and Eigen decompositions.

---

_[View the full topic](https://discourse.julialang.org/t/benchmark-matlab-julia-for-matrix-operations/2000)._
