# Tensor contration efficiency

**URL:** <https://discourse.julialang.org/t/tensor-contration-efficiency/52897>\
**Category:** Numerics\
**Created:** [January 5, 2021, 7:46pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897 "2021-01-05T19:46:59Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nellatur](https://avatars.discourse-cdn.com/v4/letter/n/76d3ee/32.png) [@Nellatur](https://discourse.julialang.org/u/Nellatur)\
**Post date:** [January 5, 2021, 7:46pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897/1 "2021-01-05T19:46:59Z")

</div>

I saw some benchmark between NumPy and BLAS from C++ regarding matrix-matrix multiplication (a special case of tensor contraction)  
[https://stackoverflow.com/questions/7596612/benchmarking-python-vs-c-using-blas-and-numpy](https://stackoverflow.com/questions/7596612/benchmarking-python-vs-c-using-blas-and-numpy)  
it seems NumPy can be slower than C++ by 1-2 orders of magnitude.

I am wondering how about einsum/TensorOperations from Julia

From [Julia Micro-Benchmarks](https://julialang.org/benchmarks/)  
it seems matrix multiplication C and python is similar, which is quite different than the above StackOverflow link. Hence, I would like to see a more comprehensive benchmark, especially comparing with using BLAS directly from C/C++/Fortran and Julia 🙂

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [January 5, 2021, 9:06pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897/2 "2021-01-05T21:06:26Z")

</div>

Both Julia and numpy call BLAS, so as long as they use the same Blas library, you should not se any difference. Numpy might use MKL by default whereas Julia defaults to openblas (MKL.jl is available though).

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [January 5, 2021, 9:07pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897/3 "2021-01-05T21:07:42Z")

</div>

For a non-BLAS based but highly competitive library, see [Tullio.jl](https://github.com/mcabbott/Tullio.jl)

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [January 5, 2021, 9:12pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897/4 "2021-01-05T21:12:13Z")

</div>

The Readme of [https://github.com/chriselrod/PaddedMatrices.jl](https://github.com/chriselrod/PaddedMatrices.jl) has some benchmarks of different Blas-like libraries for matrix multiplication.

---

<div class="post-metadata">

**Author:** ![Nellatur](https://avatars.discourse-cdn.com/v4/letter/n/76d3ee/32.png) [@Nellatur](https://discourse.julialang.org/u/Nellatur)\
**Post date:** [January 6, 2021, 7:35pm UTC](https://discourse.julialang.org/t/tensor-contration-efficiency/52897/5 "2021-01-06T19:35:21Z")

</div>

Thanks. I think NumPy used other BLAS as default since there are ways to incorporate MKL  
[https://software.intel.com/content/www/us/en/develop/articles/build-numpy-with-mkl-and-icc.html](https://software.intel.com/content/www/us/en/develop/articles/build-numpy-with-mkl-and-icc.html)

In my limited experience, NumPy-MKL/opt-einsum with optimize=True still a factor of 3~10 slower comparing with C++ loaded MKL-BLAS. Maybe I compare small scales, 1e-2~1e-5 s, contraction time cases.
