# Specialized matrix-matrix multiplication algorithm

**URL:** <https://discourse.julialang.org/t/specialized-matrix-matrix-multiplication-algorithm/116282>\
**Category:** New to Julia\
**Tags:** question, performance, linearalgebra\
**Created:** [June 27, 2024, 1:46am UTC](https://discourse.julialang.org/t/specialized-matrix-matrix-multiplication-algorithm/116282 "2024-06-27T01:46:13Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [July 9, 2024, 2:02am UTC](https://discourse.julialang.org/t/specialized-matrix-matrix-multiplication-algorithm/116282/4 "2024-07-09T02:02:26Z")

</div>

See the following discussion on what’s involved in implementing a fast matrix–matrix multiplication routine in pure Julia code (or any compiled language, for that matter), along with some example code: [Julia matrix-multiplication performance](https://discourse.julialang.org/t/julia-matrix-multiplication-performance/55175)

Also these course notes: [18335/notes/Memory-and-Matrices-latest.ipynb at spring21 · mitmath/18335 · GitHub](https://github.com/mitmath/18335/blob/spring21/notes/Memory-and-Matrices-latest.ipynb)

(A fundamental problem with implementing matrix–matrix multiplication as a sequence of matrix–vector products is that it has poor temporal memory locality, so it won’t take full advantage of the caches and you will end up being memory-bound. To do better, you have to divide the matrices into submatrix blocks rather than simply into columns; there are a variety of strategies for this, as discussed in the link above. But to get the last factor of 2–3 in performance, not including multi-threading, is pretty hard; you have to heavily optimize the low-level kernels.)

---

_[View the full topic](https://discourse.julialang.org/t/specialized-matrix-matrix-multiplication-algorithm/116282)._
