# Unexpected poor FOR loops performance

**URL:** <https://discourse.julialang.org/t/unexpected-poor-for-loops-performance/130276>\
**Category:** Performance\
**Tags:** question\
**Created:** [June 27, 2025, 11:33am UTC](https://discourse.julialang.org/t/unexpected-poor-for-loops-performance/130276 "2025-06-27T11:33:11Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 27, 2025, 1:20pm UTC](https://discourse.julialang.org/t/unexpected-poor-for-loops-performance/130276/6 "2025-06-27T13:20:44Z")

</div>

> [@andreasvarga](#):
>
> For me the important question is: which version of code should I use in an iterative solver (e.g., for solving Lyapunov matrix equations): the faster one with allocation of an n x n array, or the slower one, which performs no allocation?

If you’re doing O(n^3) work, and n is sufficiently large, it’s almost always worth the O(n^2) memory allocation to call a fast BLAS routine. That being said, if you are doing this over and over again in an iterative solver, you can probably pre-allocate any arrays you need for BLAS/LAPACK.

PS. Note that this has nothing to do with `for` loops or Julia, and everything to do with the tricks that are required to make matrix–matrix multiplications fast. These kinds of optimizations can be done in Julia too (e.g. see Octavian.jl), but require a different algorithm than textbook-style 3-nested loops: [Julia matrix-multiplication performance](https://discourse.julialang.org/t/julia-matrix-multiplication-performance/55175)

---

_[View the full topic](https://discourse.julialang.org/t/unexpected-poor-for-loops-performance/130276)._
