# LoopVec, Tullio losing to Matrix multiplication

**URL:** <https://discourse.julialang.org/t/loopvec-tullio-losing-to-matrix-multiplication/115798>\
**Category:** Performance\
**Created:** [June 18, 2024, 9:05am UTC](https://discourse.julialang.org/t/loopvec-tullio-losing-to-matrix-multiplication/115798 "2024-06-18T09:05:01Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 18, 2024, 2:54pm UTC](https://discourse.julialang.org/t/loopvec-tullio-losing-to-matrix-multiplication/115798/6 "2024-06-18T14:54:42Z")

</div>

> [@mikmoore](#):
>
> Matrix multiplication has been optimized continuously over decades, and sometimes goes so far as to feature hand-written assembly and architecture-specific implementations to maximize performance.

I would also add that optimizing matrix multiplication involves “nonlocal” changes to the code — it’s not simply a matter of taking the naive 3-nested-loop algorithm and vectorizing/multithreading/fine-tuning the loops. The whole structure of the code is changed, e.g. to “block” the algorithm to improve cache performance (with multiple nested levels of “blocking” for multiple levels of the cache, even at the lowest level treating the registers as a kind of ideal cache). This is not the sort of transformation that something like LoopVectorization.jl does.

See also this thread: [Julia matrix-multiplication performance](https://discourse.julialang.org/t/julia-matrix-multiplication-performance/55175)

---

_[View the full topic](https://discourse.julialang.org/t/loopvec-tullio-losing-to-matrix-multiplication/115798)._
