# Sparse matrix multiplication for Metal

**URL:** <https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088>\
**Category:** GPU\
**Tags:** gpu, sparse, metaljl, sparsearrays\
**Created:** [July 27, 2025, 10:06am UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088 "2025-07-27T10:06:45Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 27, 2025, 10:06am UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/1 "2025-07-27T10:06:45Z")

</div>

I need to perform a multiplication between two fairly large **sparse** matrices `C = A * B` and I would like to try using the GPU for that. The problem arises in a context where `A` represents a fixed 2D convolution kernel, while `B` represents a _list_ of 2D images (each of which is a column of `B`). Therefore, if necessary, I can easily change the storage scheme of `A` without a significant penalty.

I can easily do this operation with CUDA, but I could not find any way to perform sparse matrix products with Metal. Assuming there is no official implementation available, does anyone know a generic implementation of generic sparse matrix multiplication SpGEMM for GPUs? Ideally, one could perhaps write a suitable kernel using KernelAbstractions.

---

<div class="post-metadata">

**Author:** ![LaurentPlagne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/laurentplagne/32/10103_2.png) [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Post date:** [July 27, 2025, 1:15pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/2 "2025-07-27T13:15:55Z")

</div>

Did you try to use AppleAccelerate ?

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 27, 2025, 9:37pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/3 "2025-07-27T21:37:46Z")

</div>

@LaurentPlagne as far as I know AppleAccelerate is not using the GPU.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 27, 2025, 11:04pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/4 "2025-07-27T23:04:43Z")

</div>

For sparse arrays you can try [Sparse Arrays · The Julia Language](https://docs.julialang.org/en/v1/stdlib/SparseArrays/).

---

<div class="post-metadata">

**Author:** ![pitsianis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pitsianis/32/26588_2.png) [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Post date:** [July 28, 2025, 12:10am UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/5 "2025-07-28T00:10:00Z")

</div>

> [@Marco\_Lombardi](#):
>
> The problem arises in a context where `A` represents a fixed 2D convolution kernel, while `B` represents a _list_ of 2D images (each of which is a column of `B`).

Why don’t you try to perform the convolutions in the Fourier domain?  
You precompute the FFT of the padded kernel once and then in a loop (or stream) you multiply it element-wise with the FFT of each of the images and then take the inverse FFT of each result.

It should be faster for 5x5 and larger kernels even for a single image..

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 28, 2025, 9:49am UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/6 "2025-07-28T09:49:59Z")

</div>

My question was associated to Metal, so to the Apple GPU library. In other words, I want to use the GPU to speed up the computation.

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 28, 2025, 9:56am UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/7 "2025-07-28T09:56:31Z")

</div>

Thank you @pitsianis. The problem with this approach is that I will need to perform a matrix-matrix multiplication, not a matrix-vector multiplication. The latter could be performed using 2D FFT, but the former requires either 3D FFT or a repeated 2D FFT, and I am not sure I will gain much given the sparsity of the matrix (but I will give a try). Additionally, I am not sure Metal.jl has FFT (probably not?).

---

<div class="post-metadata">

**Author:** ![simsurace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simsurace/32/30216_2.png) [@simsurace](https://discourse.julialang.org/u/simsurace)\
**Post date:** [July 28, 2025, 2:31pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/8 "2025-07-28T14:31:21Z")

</div>

Maybe take a look at [this issue on Metal.jl](https://github.com/JuliaGPU/Metal.jl/issues/208).

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [July 28, 2025, 3:20pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/9 "2025-07-28T15:20:05Z")

</div>

There are some hardware-agnostic sparse matrix types and kernels by @AntoineBut in [GitHub - AntoineBut/GPUGraphs.jl: Workspace for my semester project in HPNALGS @ EPFL](https://github.com/AntoineBut/GPUGraphs.jl)

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 29, 2025, 3:41pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/10 "2025-07-29T15:41:41Z")

</div>

Just as a record, I tried FFT, but on the CPU it is significantly slower (by a factor ~30) than the sparse matrix approach. Also, it requires a lot of data transfer, so I do not think it is sensible to try this path on the GPU.

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 29, 2025, 3:45pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/11 "2025-07-29T15:45:50Z")

</div>

Thank you @gdalle. I had a look and could only find matrix-vector kernels, but perhaps I am missing something.

---

<div class="post-metadata">

**Author:** ![rveltz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rveltz/32/2707_2.png) [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Post date:** [July 29, 2025, 3:48pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/12 "2025-07-29T15:48:34Z")

</div>

maybe mlx has it. Id say MLX.jl does not have it yet though

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [July 29, 2025, 4:16pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/13 "2025-07-29T16:16:48Z")

</div>

Oh right, I had missed the matmat aspect. Perhaps these basic matvec kernels can be a starting point though

---

<div class="post-metadata">

**Author:** ![pitsianis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pitsianis/32/26588_2.png) [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Post date:** [July 29, 2025, 5:22pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/14 "2025-07-29T17:22:13Z")

</div>

> [@Marco\_Lombardi](#):
>
> Just as a record, I tried FFT, but on the CPU it is significantly slower (by a factor ~30) than the sparse matrix approach. Also, it requires a lot of data transfer, so I do not think it is sensible to try this path on the GPU.

I am surprised by your comment. We are using the Fourier domain approach in FastLocalCorrelationCoefficients.jl

Also, regarding the data transfers, you can work with sub images (at an added computational cost of redoing the sub-image border).

If you do not mind, please share the code of your experiment, at least one of us is going to learn something new 🙂

---

<div class="post-metadata">

**Author:** ![Marco\_Lombardi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_lombardi/32/14608_2.png) [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Post date:** [July 31, 2025, 1:29pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/15 "2025-07-31T13:29:07Z")

</div>

Thank you. I checked MLX and it has matrix multiplication. However, I did not find any indication that it has a sparse matrix type/ do you know how to use sparse matrices with MLX?

---

<div class="post-metadata">

**Author:** ![rveltz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rveltz/32/2707_2.png) [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Post date:** [July 31, 2025, 5:59pm UTC](https://discourse.julialang.org/t/sparse-matrix-multiplication-for-metal/131088/16 "2025-07-31T17:59:44Z")

</div>

nope. I am not on my computer but mlx is where I would look if metal.jl does not have what I want
