# Optimization Based on Intel MKL Matrix Multiplication Batch Mode

**URL:** <https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989>\
**Category:** Internals & Design\
**Tags:** proposal\
**Created:** [June 27, 2018, 6:15am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989 "2018-06-27T06:15:31Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![RoyiAvital](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/royiavital/32/571_2.png) [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Post date:** [June 27, 2018, 6:15am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989/1 "2018-06-27T06:15:31Z")

</div>

Intel MKL has a batch mode for Matrix Multiplication (See [Introducing Batch GEMM Operations](https://software.intel.com/en-us/articles/introducing-batch-gemm-operations)).

It seems it has great potential on speeding up small matrix multiplication operations in 2 ways (I can see, probably more):

1. Broadcasting  
When multiplying 2 2D array with 3D array and broadcasting the Matrix Multiplication operation along the 3rd dimension.
2. Lazy Evaluation  
When many small matrix operation are called in a loop and then they are accumulated and sent as a batch job for the batch mode in MKL.

Is this implemented?  
If not, could it be added to 1.x (Of course not 1.0, but on its optimization phase once it is released)?

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [June 27, 2018, 7:13am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989/2 "2018-06-27T07:13:22Z")

</div>

This seems to be basically a thin wrapper around threading. It should be supported by good support for easy threading in Julia rather than at the level of MKL.

---

<div class="post-metadata">

**Author:** ![RoyiAvital](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/royiavital/32/571_2.png) [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Post date:** [June 27, 2018, 7:51am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989/3 "2018-06-27T07:51:22Z")

</div>

@antoine-levitt,  
Indeed if the whole magic is via Multi Threading it should be don in Julia level.  
Though it might be less overhead to call one C function instead of multi calls, no?

Unless small matrices will be treated by JuliaBLAS and then engine which will Multi Threaded those operations as above will be the best choice.

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [June 27, 2018, 7:53am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989/4 "2018-06-27T07:53:51Z")

</div>

I don’t think the overhead of calling C is significant. And if you’re doing very small matrices, you’re probably better off with StaticArrays anyway, which bypasses BLAS.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [June 27, 2018, 7:59am UTC](https://discourse.julialang.org/t/optimization-based-on-intel-mkl-matrix-multiplication-batch-mode/11989/5 "2018-06-27T07:59:44Z")

</div>

TensorOperations.jl maps some tensor operations to BLAS calls already, so you can take advantage of MKL.  
[https://github.com/Jutho/TensorOperations.jl/blob/6bb64d5c83a890b2ac4e4eddb35af36d0bedf89b/src/implementation/stridedarray.jl](https://github.com/Jutho/TensorOperations.jl/blob/6bb64d5c83a890b2ac4e4eddb35af36d0bedf89b/src/implementation/stridedarray.jl)
