# v?Mul in MKL

**URL:** <https://discourse.julialang.org/t/v-mul-in-mkl/71039>\
**Category:** New to Julia\
**Tags:** performance\
**Created:** [November 6, 2021, 5:06am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039 "2021-11-06T05:06:13Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![andrew-saydjari](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrew-saydjari/32/19896_2.png) [@andrew-saydjari](https://discourse.julialang.org/u/andrew-saydjari)\
**Post date:** [November 6, 2021, 5:06am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/1 "2021-11-06T05:06:13Z")

</div>

In trying to speed up `.*`, I was wondering if much progress had been made at making v?Mul accessible (maybe as part of MKL.jl). Mostly just wondering if it is worth upgrading to 1.7 to make the install work. Currently the fastest `.*` that I can see is just to do:

```julia
function vMul_test!(A,B,C)
    n = length(A)
    @simd for i = 1 : n
        @inbounds A[i] = B[i] * C[i]
    end
end

```

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [November 6, 2021, 7:04am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/2 "2021-11-06T07:04:28Z")

</div>

Have you tried `ccall`ing into MKL and benchmarked the performance of the MKL functions? I’d be curious.

I would assume that the MKL variants are multithreaded. So you should probably also use threads in your Julia implementation. Unfortunately, we don’t have multithreaded broadcasting (yet). Maybe there exist implementations in packages?

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [November 6, 2021, 7:05am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/3 "2021-11-06T07:05:35Z")

</div>

Oh, and using MKL.jl with Julia 1.7 is a dream. Definitely worth the upgrade! 🙂

---

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [November 6, 2021, 7:19am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/4 "2021-11-06T07:19:02Z")

</div>

[FastBroadcast.jl](https://github.com/YingboMa/FastBroadcast.jl) has multithreaded broadcast for such a case.

---

<div class="post-metadata">

**Author:** ![N5N3](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/n5n3/32/17663_2.png) [@N5N3](https://discourse.julialang.org/u/N5N3)\
**Post date:** [November 6, 2021, 11:29am UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/5 "2021-11-06T11:29:34Z")

</div>

You can access MKL’s `v?Mul` via [IntelVectorMath.jl](https://github.com/JuliaMath/IntelVectorMath.jl)  
Previous bench shows that the performance of `v?Mul` is equivalent to Base or slower. So the offical release didn’t add related routines.  
You can add

```julia
def_binary_op(Float64, Float64, :multiply, :multiply!, :Mul, false)
def_binary_op(Float32, Float32, :multiply, :multiply!, :Mul, false)
def_binary_op(ComplexF64, ComplexF64, :multiply, :multiply!, :Mul, false)
def_binary_op(ComplexF32, ComplexF32, :multiply, :multiply!, :Mul, false)

```

to `src\IntelVectorMath.jl`, and call `IVM.multiply(!)` for usage.

---

<div class="post-metadata">

**Author:** ![andrew-saydjari](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrew-saydjari/32/19896_2.png) [@andrew-saydjari](https://discourse.julialang.org/u/andrew-saydjari)\
**Post date:** [November 6, 2021, 5:17pm UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/6 "2021-11-06T17:17:22Z")

</div>

Thanks for pointing out how to add the `v?Mul` to IntelVectorMath.jl (which is a really nice package!) and providing context as to why it was not added to the official release. I will check out some of the other multithreading suggestions above for an attempt at a speed up.

---

<div class="post-metadata">

**Author:** ![andrew-saydjari](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrew-saydjari/32/19896_2.png) [@andrew-saydjari](https://discourse.julialang.org/u/andrew-saydjari)\
**Post date:** [November 6, 2021, 5:51pm UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/7 "2021-11-06T17:51:13Z")

</div>

For context timings on my machine with some of the easiest multithreading solutions.

```julia
using Einsum, BenchmarkTools, LoopVectorization

A = rand(1000,1000)
B = rand(1000,1000)
C = rand(1000,1000);

function f1!(A,B,C)
    @einsum A[i,j] = B[i,j] * C[i,j]
end

function f2!(A,B,C)
    n = length(A)
    @simd for i = 1 : n
        @inbounds A[i] = B[i] * C[i]
    end
end

function f3!(A,B,C)
    n = length(A)
    @avxt for i = 1 : n
        @inbounds A[i] = B[i] * C[i]
    end
end

function f4!(A,B,C)
    @vielsum A[i,j] = B[i,j] * C[i,j]
end

```

On a single thread I find ~ 1 ms from all of the methods with `f2!` being optimal (since it has no overhead asking about if there are more threads. When using 2 threads, `f3!` seems optimal, timings below.

```julia
julia> @btime f1!($A,$B,$C);
  1.174 ms (0 allocations: 0 bytes)

julia> @btime f2!($A,$B,$C);
  1.029 ms (0 allocations: 0 bytes)

julia> @btime f3!($A,$B,$C);
  457.723 μs (0 allocations: 0 bytes)

julia> @btime f4!($A,$B,$C);
  567.861 μs (11 allocations: 1.55 KiB)

```

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [November 6, 2021, 6:19pm UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/8 "2021-11-06T18:19:03Z")

</div>

I wouldn’t expect MKL to be faster than Julia on plain broadcasted multiplication. Is there any reason to?

> [@andrew-saydjari](#):
>
> `@avxt for i = 1 : n`

Are you using an old version of LoopVectorization.jl? Or are the `@avx`/`@avxt` still available along with the new `@turbo`/`@tturbo`?

---

<div class="post-metadata">

**Author:** ![andrew-saydjari](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrew-saydjari/32/19896_2.png) [@andrew-saydjari](https://discourse.julialang.org/u/andrew-saydjari)\
**Post date:** [November 6, 2021, 6:26pm UTC](https://discourse.julialang.org/t/v-mul-in-mkl/71039/9 "2021-11-06T18:26:12Z")

</div>

Whoops, I was looking at old docs, but using a current version (v0.12.66), It appears both are still available, but thanks for the comment. I will switch to `@tturbo`.

And, I guess I don’t have a good enough intuition to know whether or not to expect the MKL solution would be faster which is why I wanted to just do the experiment.
