# How to improve the speed of matrix multiplication?

**URL:** <https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293>\
**Category:** Performance\
**Tags:** question\
**Created:** [January 13, 2021, 5:20pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293 "2021-01-13T17:20:27Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![zxjroger](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zxjroger/32/11826_2.png) [@zxjroger](https://discourse.julialang.org/u/zxjroger)\
**Post date:** [January 13, 2021, 5:20pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/1 "2021-01-13T17:20:27Z")

</div>

I have a code about large matrix multiplication in a for loop which currently takes longer time than I expected. I wonder if it is possible to improve the following MWE code:

```julia
temp1 = randn(1600, 1600)
temp2 = randn(1600)
temp = Vector{Vector{Float64}}(undef, 10000)
@time for t = 1:10000
        temp[t] = temp1 * temp2
end

6.276464 seconds (19.49 k allocations: 123.436 MiB, 0.34% gc time)

```

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [January 13, 2021, 5:20pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/2 "2021-01-13T17:20:43Z")

</div>

Use a GPU

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [January 13, 2021, 5:27pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/3 "2021-01-13T17:27:52Z")

</div>

> [@zxjroger](#):
>
> longer time than I expected

you expected wrong. Directly multiplying float64 matrices simply calls OpenBLAS under the hood, there’s nothing you can do (in fact, probably nothing can be done period)

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [January 13, 2021, 5:41pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/4 "2021-01-13T17:41:23Z")

</div>

> [@jling](#):
>
> you expected wrong. Directly multiplying float64 matrices simply calls OpenBLAS under the hood, there’s nothing you can do

Not sure that’s totally fair; for example, you can try changing the number of threads OpenBLAS uses, because apparently its defaults are sometimes pretty bad. You can also try MKL.jl to use MKL instead.

---

<div class="post-metadata">

**Author:** ![fgerick](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fgerick/32/13228_2.png) [@fgerick](https://discourse.julialang.org/u/fgerick)\
**Post date:** [January 13, 2021, 5:47pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/5 "2021-01-13T17:47:27Z")

</div>

Not much faster for this example, but you can avoid some allocations by using

```julia
@time for t = 1:10000
    mul!(temp[t],temp1,temp2)
 end

```

but you have to allocate them properly before…

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 13, 2021, 6:32pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/6 "2021-01-13T18:32:21Z")

</div>

Especially for matrix-vector multiplication, which is totally limited by memory bandwidth. When using a single thread, `temp1 * temp2` and `sum(temp1)` will be about the same fast:

```julia
julia> temp1 = randn(1600, 1600);

julia> temp2 = randn(1600);

julia> temp3 = similar(temp2);

julia> using LinearAlgebra; BLAS.set_num_threads(1);

julia> @benchmark mul!($temp3, $temp1, $temp2)
BenchmarkTools.Trial:
  memory estimate: 0 bytes
  allocs estimate: 0
  --------------
  minimum time: 1.074 ms (0.00% GC)
  median time: 1.099 ms (0.00% GC)
  mean time: 1.115 ms (0.00% GC)
  maximum time: 1.499 ms (0.00% GC)
  --------------
  samples: 4479
  evals/sample: 1

julia> @benchmark sum($temp1)
BenchmarkTools.Trial:
  memory estimate: 0 bytes
  allocs estimate: 0
  --------------
  minimum time: 1.120 ms (0.00% GC)
  median time: 1.149 ms (0.00% GC)
  mean time: 1.169 ms (0.00% GC)
  maximum time: 1.613 ms (0.00% GC)
  --------------
  samples: 4274
  evals/sample: 1

```

(Re: the GPU suggestion, GPUs do have far higher memory bandwidth than any normal CPU, so they are good for this sort of memory bandwidth limited thing [and not just compute heavy things like matrix-matrix multiplication])

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [January 13, 2021, 6:51pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/7 "2021-01-13T18:51:29Z")

</div>

You are benchmarking in global scope, which is not a good idea.

Also, if you have some memory you can reuse, you can write

```julia
mul!(out, temp1, temp2)

```

instead.

And if your matrices have some sort of special structure, that may be exploited.

---

<div class="post-metadata">

**Author:** ![moeddel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moeddel/32/18641_2.png) [@moeddel](https://discourse.julialang.org/u/moeddel)\
**Post date:** [January 13, 2021, 9:44pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/8 "2021-01-13T21:44:25Z")

</div>

Are your consecutive iterations dependent of each other? If not, you could maybe reformulate your problem to use matrix-matrix multiplication, which is not memory bound and way more efficient.

```julia
using BenchmarkTools
temp1 = randn(1600, 1600)
temp2 = randn(1600)

@benchmark $temp1 * $temp2
  814.858 μs (1 allocation: 12.63 KiB)

```

would take more than 8.1 s per 10000 iterations, while the same calculation reformulated to a matrix-matrix problem

```julia
temp3 = repeat($temp2, 1, 10000)
@btime $temp1 * $temp3
  952.054 ms (2 allocations: 122.07 MiB)

```

is about 8 times faster on my machine.

---

<div class="post-metadata">

**Author:** ![moeddel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moeddel/32/18641_2.png) [@moeddel](https://discourse.julialang.org/u/moeddel)\
**Post date:** [January 14, 2021, 8:26am UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/9 "2021-01-14T08:26:25Z")

</div>

If your application allows for it you could also switch to single precision `Float32`, which roughly reduces the time by a factor of two if you compare with the benchmark above

```julia
using BenchmarkTools

temp1 = randn(Float32,1600, 1600)
temp2 = randn(Float32,1600)

@btime $temp1 * $temp2
  518.491 μs (1 allocation: 6.38 KiB)

```

---

<div class="post-metadata">

**Author:** ![zxjroger](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zxjroger/32/11826_2.png) [@zxjroger](https://discourse.julialang.org/u/zxjroger)\
**Post date:** [January 14, 2021, 4:12pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/10 "2021-01-14T16:12:36Z")

</div>

I really appreciate your suggestion. I am able to reduce the time from 10 seconds to 2 seconds. Thank you very much.

---

<div class="post-metadata">

**Author:** ![tamasgal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamasgal/32/27946_2.png) [@tamasgal](https://discourse.julialang.org/u/tamasgal)\
**Post date:** [January 14, 2021, 4:34pm UTC](https://discourse.julialang.org/t/how-to-improve-the-speed-of-matrix-multiplication/53293/11 "2021-01-14T16:34:39Z")

</div>

The hint from @DNF is valuable and you should think about it. Maybe you can even improve the whole approach or do something smart when your matrices are somehow correlated or have a specific structure or property (something like symmetric, logarithmic value range, etc.).
