# MKLSparse with AMD cpu

**URL:** <https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559>\
**Category:** Performance\
**Created:** [November 22, 2020, 12:01am UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559 "2020-11-22T00:01:07Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bruno\_Amorim](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@Bruno\_Amorim](https://discourse.julialang.org/u/Bruno_Amorim)\
**Post date:** [November 22, 2020, 12:01am UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/1 "2020-11-22T00:01:07Z")

</div>

As discussed in other threads

- [Acceleration of Intel MKL on AMD Ryzen CPU’s](https://discourse.julialang.org/t/acceleration-of-intel-mkl-on-amd-ryzen-cpus/31287)
- [Hack: AMD Ryzen/TR/Epyc + Intel Math Kernel Library (MKL)](https://discourse.julialang.org/t/hack-amd-ryzen-tr-epyc-intel-math-kernel-library-mkl/31226)

Intel MKL descriminates against AMD cpu’s (although this might be changing).

Does this affect the performance of MKLSparse.jl on AMD cpu’s? Does anyone have experinece in this?

Ideally, one would have a AOCLSparse.jl to use [AMD’s AOCL](https://github.com/amd/aocl-sparse) for amd cpu’s. But that doesn’t seem to exist yet.

---

<div class="post-metadata">

**Author:** ![pablosanjose](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pablosanjose/32/7006_2.png) [@pablosanjose](https://discourse.julialang.org/u/pablosanjose)\
**Post date:** [November 23, 2020, 12:41pm UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/2 "2020-11-23T12:41:21Z")

</div>

I’m also very interested in this. In particular, can AMD chips (I’m thinking Zen3 in particular) multithread sparse mat-vec multiplication efficiently using MKL?

---

<div class="post-metadata">

**Author:** ![jamblejoe](https://avatars.discourse-cdn.com/v4/letter/j/ee7513/32.png) [@jamblejoe](https://discourse.julialang.org/u/jamblejoe)\
**Post date:** [October 14, 2021, 8:23pm UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/3 "2021-10-14T20:23:15Z")

</div>

I am also very interested in that. For the exact same reason as @pablosanjose . Any updates?

---

<div class="post-metadata">

**Author:** ![Bruno\_Amorim](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@Bruno\_Amorim](https://discourse.julialang.org/u/Bruno_Amorim)\
**Post date:** [October 15, 2021, 10:25am UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/4 "2021-10-15T10:25:54Z")

</div>

I have an update on this. I got a desktop with a Ryzen 7 4750G cpu (Zen 2, 8 cores, 3.6-4.4GHz, 8MB L3 cache). Appatenly MKLSparse does work with AMD cpus and without any hacking now.

By simply doing  
`pkg> add MKLSpare`

and then testing:

```julia
using SparseArrays, BenchmarkTools, LinearAlgebra
A = sprand(10_000, 10_000, 0.001)
v = rand(10_000)

```

Without MKLSparse I got the timings:

```julia
@btime A*v;
  105.259 μs (2 allocations: 78.20 KiB)

@btime A'v;
  97.819 μs (3 allocations: 78.25 KiB)

```

With MKLSparse:

```julia
using MKLSparse
@btime A*v;
  62.430 μs (2 allocations: 78.20 KiB)
@btime A'v;
  13.730 μs (3 allocations: 78.25 KiB)

```

MKLSparse is using the 8 cores of the cpu.

I also tested against `ThreadedSparseArrays.jl`

```julia
using ThreadedSparseArrays
tA = ThreadedSparseMatrixCSC(A);
@btime tA*v;
  126.399 μs (20 allocations: 79.25 KiB)
@btime tA'v;
  22.820 μs (63 allocations: 83.20 KiB)

```

Here the mat-vec multiplication does not appear to be using threads, while the adjoint(mat)-vec did use the 8 cores (but MKLSparse still beats it).

---

<div class="post-metadata">

**Author:** ![Bruno\_Amorim](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@Bruno\_Amorim](https://discourse.julialang.org/u/Bruno_Amorim)\
**Post date:** [October 15, 2021, 10:36am UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/5 "2021-10-15T10:36:24Z")

</div>

Interesting… the time difference between`ThreadedSparseArrays.jl` vs `MKLSparse.jl` for the adjoint(mat)-vec goes away if we consider larger matrices:

```julia
using SparseArrays, BenchmarkTools

A = sprand(100_000, 100_000, 0.001)
v = rand(100_000)

using ThreadedSparseArrays
tA = ThreadedSparseMatrixCSC(A);
@btime tA'v;
  5.495 ms (64 allocations: 786.33 KiB)

using MKLSparse
@btime A'v;
  5.449 ms (3 allocations: 781.38 KiB)

```

So probably the slowness of `ThreadedSparseArrays` in the previous example is due to the overhead of Julia threads?

I wonder if `Polyester.jl` could improve things

---

<div class="post-metadata">

**Author:** ![freestatelabs](https://avatars.discourse-cdn.com/v4/letter/f/9de0a6/32.png) [@freestatelabs](https://discourse.julialang.org/u/freestatelabs)\
**Post date:** [December 11, 2024, 8:25pm UTC](https://discourse.julialang.org/t/mklsparse-with-amd-cpu/50559/6 "2024-12-11T20:25:18Z")

</div>

I’ve been looking into this recently (I wanted to experiment with AOCL vs MKL benchmarks on my AMD PC) - has there been any change since you originally posted this thread in 2020?
