# Acceleration of Intel MKL on AMD Ryzen CPU's

**URL:** <https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287>\
**Category:** Performance\
**Tags:** performance, mkl, linearalgebra\
**Created:** [November 19, 2019, 9:35pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287 "2019-11-19T21:35:18Z")\
**Posts on this page:** 15\
**Page:** 2

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 8, 2024, 4:33pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/21 "2024-04-08T16:33:49Z")

</div>

The `svd()` function of LinearAlgebra works on variables that are `StaticArrays` or `StaticVectors`. To which degrees this is speed optimized I don’t know, try it yourself…

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 8, 2024, 4:34pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/22 "2024-04-08T16:34:55Z")

</div>

So I can use StaticArrays but not LoopVectorization?

---

<div class="post-metadata">

**Author:** ![maxfreu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxfreu/32/17468_2.png) [@maxfreu](https://discourse.julialang.org/u/maxfreu)\
**Post date:** [April 9, 2024, 8:20pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/23 "2024-04-09T20:20:13Z")

</div>

Yes, you can use them independently.

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 10, 2024, 2:48pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/24 "2024-04-10T14:48:34Z")

</div>

Please, correct me if I am mistaken:  
Well, it seems that the answer is that svd from LinearAlgebra converts the input matrices into static matrices, and then passes them to the svd routine of LAPACK. So it is taking advantage of static matrices.  
Regarding the LoopVectorization, svd from LinearAlgebra does not use it. But LinearAlgebra uses the ScaLAPACK, which is multi-threaded, and hence, has the same advantage than LoopVectorization.  
On the other hand, none of this can be changed if I put

```julia
using StaticArrays or using LoopVectorization
using LinearAlgebra

```

id est, the behavior of LinearAlgebra will not change if I import any of StaticArrays or LoopVectorization before importing LinearAlgebra.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 10, 2024, 2:59pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/25 "2024-04-10T14:59:12Z")

</div>

> [@aasdelat](#):
>
> which is multi-threaded, and hence, has the same advantage than LoopVectorization

This is a misunderstanding. The main advantage of LoopVectorization is that it is using SIMD instructions (single instruction, multiple data) very effectively, which is orthogonal to multithreading.

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 10, 2024, 3:01pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/26 "2024-04-10T15:01:00Z")

</div>

Ok, but, does ScaLAPCK also use SIMD?, does LoopVectorization also use multi-threading?

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 10, 2024, 3:04pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/27 "2024-04-10T15:04:17Z")

</div>

> [@aasdelat](#):
>
> does LoopVectorization also use multi-threading

It can also use multi-threading, if that is what you want:

[https://juliasimd.github.io/LoopVectorization.jl/stable/examples/multithreading/#Multithreading](https://juliasimd.github.io/LoopVectorization.jl/stable/examples/multithreading/#Multithreading)

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 10, 2024, 3:25pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/28 "2024-04-10T15:25:58Z")

</div>

Ok, I understand now.  
But the help on LoopVectorization is not clear about the purpose of this library. I think this is a general issue about help on any Julia package. I think that the first things that the help on LoopVectorization should say are that it purpose is to provide matrix operations that use SIMD and multi-threading, and also in which libraries (ScaLAPACK?) it is based in order to accomplish those purposes.  
My question about the use of SIMD by ScaLAPCK, still remains.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 10, 2024, 3:51pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/29 "2024-04-10T15:51:40Z")

</div>

> [@aasdelat](#):
>
> My question about the use of SIMD by ScaLAPCK, still remains.

I’m not familiar with ScaLAPACK, but any half-decent normal LAPACK library uses SIMD. Reference BLAS + Netlib might not.  
I would assume the same for ScaLAPACK.

> [@aasdelat](#):
>
> Well, it seems that the answer is that svd from LinearAlgebra converts the input matrices into static matrices, and then passes them to the svd routine of LAPACK. So it is taking advantage of static matrices.

No, they don’t convert to static matrices.

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 11, 2024, 10:48am UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/30 "2024-04-11T10:48:12Z")

</div>

> [@Elrod](#):
>
> > [@aasdelat](#):
> >
> > So it is taking advantage of static matrices.
> 
> No, they don’t convert to static matrices.

So, svd from LinearAlgebra, does not take advantage from static matrices, but it seems to take advantage from StridedMatrix. It converts the input matrix into a StridedMatrix before passing it to the LAPACK.gesvd rouotine, that makes the svd. And I think the goals of StridedMatrix are similar to that of static matrices.

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 12, 2024, 9:58pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/31 "2024-04-12T21:58:17Z")

</div>

I want to use one aocl routine to accelerate svd on Ryzen, but this seems not to be the topic for this thread, so I have opened another one. For the interested ones:

> [@AOCL (not MKL) AMD acceleration on Ryzen CPU's](https://discourse.julialang.org/t/aocl-not-mkl-amd-acceleration-on-ryzen-cpus/112890):
>
> Hi: I am interested in calling a lapack routine, dgesvd, that makes a singular value decomposition of a matrix. The library LinearAlgebra.jl already does it, but it is a generic library that does not take full advantage of the concrete CPU you use. For Intel CPU’s, there is MKL, that can be used in Julia by means of the MKL.jl library. In order to use it, you simply import these libraries in the following order: using MKL using LinearAlgebra In this way, the library LinearAlgebra will call …

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [April 25, 2024, 9:44am UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/32 "2024-04-25T09:44:06Z")

</div>

> [@aasdelat](#):
>
> > [@Elrod](#):
> >
> > > [@aasdelat](#):
> > >
> > > So it is taking advantage of static matrices.
> > 
> > No, they don’t convert to static matrices.
> 
> So, svd from LinearAlgebra, does not take advantage from static matrices, but it seems to take advantage from StridedMatrix. It converts the input matrix into a StridedMatrix before passing it to the LAPACK.gesvd rouotine, that makes the svd. And I think the goals of StridedMatrix are similar to that of static matrices.

It converts to strided matrices, but does not use the StridedMatrix library. Julia Base library implements this type of matrix in julia/base/base/reinterpretarray.jl  
This type is based on DenseMatrices and FastContiguousSubArray. I know what DenseMatrices are, but not what FastContiguousSubArray are.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 26, 2024, 6:18pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/33 "2024-04-26T18:18:14Z")

</div>

It doesn’t convert to `StridedArray`; it works on things that already are.

It is more general than `FastContiguousSubArray`. `FastContiguousSubArray` is about supporting linear indexing:

```julia
julia> A = rand(5,5);

julia> view(A, :, 1:4) isa Base.FastContiguousSubArray
true

julia> view(A, 1:5, 1:4) isa Base.FastContiguousSubArray
false

julia> Base.IndexStyle(typeof(view(A, :, 1:4)))
IndexLinear()

julia> Base.IndexStyle(typeof(view(A, 1:4, 1:4)))
IndexCartesian()

```

That is, there is no gap between columns of a `FastContiguousSubArray`, so you can efficiently iterate over all elements without needing nested loops or cartesian indices; `eachindex`, for example, returns just `OneTo(length(A))`:

```julia
julia> eachindex(view(A, :, 1:4))
Base.OneTo(20)

julia> eachindex(view(A, 1:4, 1:4))
CartesianIndices((4, 4))

```

Many BLAS and LAPACK operations are more general than that. If they need to view an array as a two-dimensional object instead of one-dimensional anyway, the cost is low for them to support `stride(A,2) != size(A,1)`.  
In fact, this can be important to support as an implementation detail, as algorithms often work blockwise, meaning to implement them, they’ll be taking views of the full matrix anyway; whether that full matrix was itself a submatrix to begin with wasn’t so important (or alternatively, you may want to pad columns to align them).

So, `StridedArray` is more general, as a subarray can still be strided, even if not contiguous:

```julia
julia> view(A, 1:4, 1:4) isa Base.FastContiguousSubArray
false

julia> view(A, 1:4, 1:4) isa Base.StridedArray
true

julia> view(A, 1:2:4, 1:4) isa Base.StridedArray
true

julia> view(A, [1,3], 1:2:4) isa Base.StridedArray
false

```

---

<div class="post-metadata">

**Author:** ![photor](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/photor/32/14343_2.png) [@photor](https://discourse.julialang.org/u/photor)\
**Post date:** [April 27, 2024, 2:21am UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/34 "2024-04-27T02:21:56Z")

</div>

Very clear explanation!

---

<div class="post-metadata">

**Author:** ![aasdelat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aasdelat/32/45268_2.png) [@aasdelat](https://discourse.julialang.org/u/aasdelat)\
**Post date:** [May 9, 2024, 12:07pm UTC](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287/35 "2024-05-09T12:07:11Z")

</div>

Thank you very much!  
Very clear!  
I now understand a lot more. I think this should be included somewhere in the documentation.

[Previous page](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287.md?page=1)
