# Why is for loops so slow?

**URL:** <https://discourse.julialang.org/t/why-is-for-loops-so-slow/99199>\
**Category:** Performance\
**Tags:** performance\
**Created:** [May 22, 2023, 7:57am UTC](https://discourse.julialang.org/t/why-is-for-loops-so-slow/99199 "2023-05-22T07:57:04Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 22, 2023, 11:29am UTC](https://discourse.julialang.org/t/why-is-for-loops-so-slow/99199/6 "2023-05-22T11:29:57Z")

</div>

> [@perebalazs](#):
>
> make a faster equation solver than the built-in code in julia

To elaborate on @mstewart’s answer above, even for the simpler problem of _multiplying_ two m \times m matrices, which naively takes about 10 lines of code (3 nested loops), the simplest implementations (in any language!) are typically _orders of magnitude_ slower than highly optimized algorithms (which do the “same” \sim 2m^3 floating-point operations; the theoretical [lower-complexity matrix-mult algorithms](https://en.wikipedia.org/wiki/Computational_complexity_of_matrix_multiplication) are virtually never used). Optimized matrix-multiply [“BLAS” libraries](https://en.wikipedia.org/wiki/Basic_Linear_Algebra_Subprograms) often devote \gtrsim 10,000 lines of code to this problem! It’s not a question of the speed of “for loops” — optimized code completely re-organizes the algorithm into “block” operations that have better [memory locality](https://en.wikipedia.org/wiki/Locality_of_reference), for example.

See also the discussion thread: [Julia matrix-multiplication performance](https://discourse.julialang.org/t/julia-matrix-multiplication-performance/55175)

It is _very_ hard to beat (or even come close to) the performance of highlyoptimized libraries for basic operations on generic dense matrices, even if you understand these performance issues! This is true in _any_ language, even in compiled languages like C. (Where you can do better is if your matrices are very special and and you can exploit that for improved algorithms.)

> [@perebalazs](#):
>
> ```julia
> function symgauss!(A::Matrix{Float64}, b::Vector{Float64})
> @assert size(A, 1) == size(A, 2) && size(A, 1) == length(b) "Size mismatch"
> # Gauss elimination
> n = length(b)
> d::Float64 = 0
> e::Float64 = 0
> 
> ```

Note, by the way, that all of these type declarations are _not required_ for performance. As long as you initialize your variables [in a type-stable way](https://docs.julialang.org/en/v1/manual/performance-tips/#Avoid-changing-the-type-of-a-variable) like `d = zero(eltype(b)); e = zero(eltype(A))`, you can omit all of the type declarations and the code will be just as fast … and more generic, since it will work on matrices of any numeric type. (This is a common misconception.) See [“Argument-Type Declarations” in the Julia manual](https://docs.julialang.org/en/v1/manual/functions/#Argument-type-declarations).

---

_[View the full topic](https://discourse.julialang.org/t/why-is-for-loops-so-slow/99199)._
