# How to write multithreading matrix multiplication

**URL:** <https://discourse.julialang.org/t/how-to-write-multithreading-matrix-multiplication/121865>\
**Category:** New to Julia\
**Tags:** parallel, multithreading\
**Created:** [October 28, 2024, 11:23am UTC](https://discourse.julialang.org/t/how-to-write-multithreading-matrix-multiplication/121865 "2024-10-28T11:23:26Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [October 28, 2024, 12:19pm UTC](https://discourse.julialang.org/t/how-to-write-multithreading-matrix-multiplication/121865/2 "2024-10-28T12:19:07Z")

</div>

Hi,

one mistake is to call `@btime multiplyMatrices_oneThreadLoop(A, B, N)` instead of `@btime multiplyMatrices_oneThreadLoop($A, $B, $N)`. This avoids issues with [global variables](https://docs.julialang.org/en/v1/manual/performance-tips/#Avoid-global-variables).  
Further, you do assignments instead of `+=` in the loop. So your code did not calculate a matmul.

This way of implementing a matmul is also the obvious one but unfortunately quite slow.

See also this comment:

> [@Julia matrix-multiplication performance](https://discourse.julialang.org/t/julia-matrix-multiplication-performance/55175/12):
>
> For example, here is a little recursive implementation of a [cache-oblivious matrix multiplication](https://en.wikipedia.org/wiki/Cache-oblivious_algorithm) that stays within a factor of 2 of the single-threaded OpenBLAS performance on my laptop up to 3000×3000 matrices: function add\_matmul\_rec!(m,n,p, i0,j0,k0, C,A,B) if m+n+p \<= 256 # base case: naive matmult for sufficiently large matrices @avx for i = 1:m, k = 1:p c = zero(eltype(C)) for j = 1:n @inbounds c += A[i0+i,j0+j] \* B[j0+j,k0+k] …

This version of the code gives me the desired results, second runs are not faster or slower. I reduced `N` to have more reasonable runtimes.

```julia
 using Base.Threads
 using BenchmarkTools
 
 function multiplyMatrices_oneThreadLoop(A::Matrix{Float64}, B::Matrix{Float64}, N::Int64)
         C = zeros(N, N)
         Threads.@threads for i in 1:N
                 for j in 1:N
                         for k in 1:N
                                 C[i, j] += A[i, k] * B[k, j]
                         end
                 end
 
         end
         return C
 end
 
 
 function multiplyMatrices_spawnExample(A::Matrix{Float64}, B::Matrix{Float64}, N::Int64)
         C = zeros(N, N)
          @sync Threads.@spawn for i in 1:N
                 for j in 1:N
                         for k in 1:N
                                 C[i, j] += A[i, k] * B[k, j]
                         end
                 end
 
         end
         return C
 end
 
 
 
 function multiplyMatrices_default(A::Matrix{Float64}, B::Matrix{Float64}, N::Int64)
         C = zeros(N,N)
         for i in 1:N
                 for j in 1:N
                         for k in 1:N
                                 C[i, j] += A[i, k] * B[k, j]
                         end
                 end
 
         end
         return C
 end
 
 N = 100
 A = rand(N, N);
 B = rand(N, N);
 println("multi-threaded loop 1st run")
 @btime multiplyMatrices_oneThreadLoop(A, B, N)
 println("using sync spawn 1st run")
 @btime multiplyMatrices_spawnExample($A,$B,$N)
 println("default multiplication 1st run")
 @btime multiplyMatrices_default($A, $B, $N)
 println("multi-threaded loop 2nd run")
 @btime multiplyMatrices_oneThreadLoop($A, $B, $N)
 println("using sync spawn 2nd run")
 @btime multiplyMatrices_spawnExample($A,$B,$N)
 println("default multiplication 2nd run")
 @btime multiplyMatrices_default($A, $B, $N)
 println("multi-threaded loop 3rd run")
 @btime multiplyMatrices_oneThreadLoop($A, $B, $N)
 println("using sync spawn 3rd run")
 @btime multiplyMatrices_spawnExample($A,$B,$N)
 println("default multiplication 3rd run")
 @btime multiplyMatrices_default($A, $B, $N)
~                                              
``
```

---

_[View the full topic](https://discourse.julialang.org/t/how-to-write-multithreading-matrix-multiplication/121865)._
