# Loops vs Vectorization? When to use?

**URL:** <https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724>\
**Category:** New to Julia\
**Tags:** performance\
**Created:** [February 16, 2023, 2:10pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724 "2023-02-16T14:10:43Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![marianoarnaiz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marianoarnaiz/32/19377_2.png) [@marianoarnaiz](https://discourse.julialang.org/u/marianoarnaiz)\
**Post date:** [February 16, 2023, 2:10pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/1 "2023-02-16T14:10:43Z")

</div>

Hi everyone.  
I am optimizing some older codes of mine. I find that sometimes vectorizing is faster than looping (which did not happen before) Could anyone explain when it is better to use one or the other in Julia.

Thanks 🙂

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [February 16, 2023, 2:21pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/2 "2023-02-16T14:21:28Z")

</div>

In most cases the advantage of a vectorized operation comes from the guarantee that the operation is inbounds (the broadcast fails if the dimensions do not match), something that you might have to inform explicitly the compiler if writing a loop.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 16, 2023, 2:23pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/3 "2023-02-16T14:23:21Z")

</div>

It can also help if things are type unstable by providing a function barrier.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 17, 2023, 12:18am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/4 "2023-02-17T00:18:47Z")

</div>

Chris, could you please decode your important statement for the noobs with a simple example?

In the code sample below, looping via a comprehension performs better than vectorization (if I understand the term correctly) but I do not see a function barrier:

```julia
A = ["abcd", [1,2,10], π]
length.(A) # 661 ns (9 allocs: 416 bytes)
[length(a) for a in A] # 242 ns (3 allocs: 112 bytes)

```

---

<div class="post-metadata">

**Author:** ![zeroexcuses](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zeroexcuses/32/46846_2.png) [@zeroexcuses](https://discourse.julialang.org/u/zeroexcuses)\
**Post date:** [February 17, 2023, 12:45am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/5 "2023-02-17T00:45:39Z")

</div>

> [@marianoarnaiz](#):
>
> I find that sometimes vectorizing is faster than looping (which did not happen before)

I’m new to Julia. Can you please explain when vectorizing would be SLOWER than looping? Intuition says vectorizing (when possible) is always faster.

---

<div class="post-metadata">

**Author:** ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)\
**Post date:** [February 17, 2023, 1:16am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/6 "2023-02-17T01:16:04Z")

</div>

> [@rafael.guerra](#):
>
> but I do not see a function barrier

I’m sure you’re aware, that array comprehensions compile to a `Generator` (with an anonymous function), which is immediately `collect`ed into an array. (see the bottom six lines)

```julia
julia> Meta.@lower [length(a) for a in A]
:($(Expr(:thunk, CodeInfo(
    @ none within `top-level scope`
1 ─ $(Expr(:thunk, CodeInfo(
    @ none within `top-level scope`
1 ─ global var"#13#14"
│ const var"#13#14"
│ %3 = Core._structtype(Main, Symbol("#13#14"), Core.svec(), Core.svec(), Core.svec(), false, 0)
│ Core._setsuper!(%3, Core.Function)
│ var"#13#14" = %3
│ Core._typebody!(%3, Core.svec())
└── return nothing
)))
│ %2 = Core.svec(var"#13#14", Core.Any)
│ %3 = Core.svec()
│ %4 = Core.svec(%2, %3, $(QuoteNode(:(#= none:0 =#))))
│ $(Expr(:method, false, :(%4), CodeInfo(
    @ none within `none`
1 ─ %1 = length(a)
└── return %1
)))
│ #13 = %new(var"#13#14")
│ %7 = #13
│ %8 = Base.Generator(%7, A)
│ %9 = Base.collect(%8)
└── return %9
))))

```

---

<div class="post-metadata">

**Author:** ![fft](https://avatars.discourse-cdn.com/v4/letter/f/e8c25b/32.png) [@fft](https://discourse.julialang.org/u/fft)\
**Post date:** [February 17, 2023, 2:47am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/7 "2023-02-17T02:47:14Z")

</div>

Vectorizing an operation is executing a loop under the hood. In theory a loop and a vectorized operation are doing the same thing.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 17, 2023, 9:02am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/8 "2023-02-17T09:02:21Z")

</div>

> [@uniment](#):
>
> I’m sure you’re aware, that array comprehensions compile to a `Generator` (with an anonymous function), which is immediately `collect`ed into an array.

Thank you and no, I am not. I only know Julia from the walks in the park, but not from seeing her under the operating table.

---

<div class="post-metadata">

**Author:** ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)\
**Post date:** [February 17, 2023, 9:14am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/9 "2023-02-17T09:14:44Z")

</div>

Ah, then you might find [this interesting](https://discourse.julialang.org/t/call-fastclosures-jl-closure-in-array-comprehensions/93622) 😉

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [February 17, 2023, 9:51am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/10 "2023-02-17T09:51:05Z")

</div>

> [@zeroexcuses](#):
>
> Intuition says vectorizing (when possible) is always faster.

Vectorized expressions sometimes create intermediate temporary arrays which are unnecessary. In many cases you can avoid this, but not always.

With a loop, you can sometimes get much better performance:

```julia
# vectorized
foo(A) = maximum(abs2.(A))

# loop
function bar(A)
    val = zero(eltype(A))
    for i in eachindex(A)
        @inbounds val = max(val, abs2(A[i]))
    end
    return val
end

```

Benchmarks:

```julia
julia> using BenchmarkTools

julia> A = rand(-100:100, 10_000);

julia> @btime foo($A);
  7.600 μs (2 allocations: 78.17 KiB)

julia> @btime bar($A);
  2.322 μs (0 allocations: 0 bytes)

```

As you see, the loop is much faster. But it’s not always easy to predict. If I supply a vector of Floats instead of a vector of Ints, the benchmarks are very different.

(Sorry, made an error in the first draft.)

---

<div class="post-metadata">

**Author:** ![mdogan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mdogan/32/47244_2.png) [@mdogan](https://discourse.julialang.org/u/mdogan)\
**Post date:** [February 17, 2023, 9:56am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/11 "2023-02-17T09:56:53Z")

</div>

You may also want to check the “performance tips” from the official doc.

[Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/)

Maybe you’ll find some good tips.

---

<div class="post-metadata">

**Author:** ![LaurentPlagne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/laurentplagne/32/10103_2.png) [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Post date:** [February 17, 2023, 10:10am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/12 "2023-02-17T10:10:42Z")

</div>

Maybe relevant for you, broadcast is also useful to build generic code running on both GPU and CPU.  
For example, considering a n\_s\times n\_s mesh, a basic finite difference scheme can be written as

```julia
function update_force!(fx::AbstractArray{Float64,2},xc::AbstractArray{Float64,2},ks)
    ns=size(xc,1)
    r0=1:ns-2
    r1=2:ns-1
    r2=3:ns

    @views @. fx[r1,r1] = -ks*(4xc[r1,r1]
                            -xc[r1,r0] # Down
                            -xc[r0,r1] # Left
                            -xc[r2,r1] # Right
                            -xc[r1,r2]) # Up
end

```

which corresponds to the loop version that will not directly run on GPUs.

```julia
function update_force!(fx::Array{Float64,2},xc::Array{Float64,2},ks)
    nsx,nsy=size(xc)
    @inbounds for j ∈ 2:nsy-1
        for i ∈ 2:nsx-1
            fx[i,j] = -ks*(4xc[i,j]
            -xc[i,j-1]
            -xc[i-1,j]
            -xc[i+1,j]
            -xc[i,j+1])
        end
    end
end
```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [February 17, 2023, 11:46am UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/13 "2023-02-17T11:46:12Z")

</div>

> [@DNF](#):
>
> `foo(A) = maximum(abs2.(A))`

Even the version that does not artificially create an intermediate,

```julia
foo(A) = maximum(abs2,A)

```

turns out to be slower than the loop:

```julia
julia> @btime foo($A);
  1.695 μs (0 allocations: 0 bytes)

julia> @btime bar($A);
  860.459 ns (0 allocations: 0 bytes)

```

(why in this case, I don’t know, but it seems that the loop version is being able to SIMD better).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 17, 2023, 6:55pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/14 "2023-02-17T18:55:00Z")

</div>

> [@rafael.guerra](#):
>
> Chris, could you please decode your important statement for the noobs with a simple example?
> 
> In the code sample below, looping via a comprehension performs better than vectorization (if I understand the term correctly) but I do not see a function barrier:
> 
> ```julia
> A = ["abcd", [1,2,10], π]
> length.(A) # 661 ns (9 allocs: 416 bytes)
> [length(a) for a in A] # 242 ns (3 allocs: 112 bytes)
> 
> ```

A function barrier helps if you’re doing a lot of work behind the barrier, but `length` is not a lot of work.

Here is a small real-world example that may come up frequently:

```julia
using DataFrames
df = DataFrame(x = rand(1024); y = rand(1024));

function loop!(df)
    (; x, y) = df
    z = similar(x)
    @inbounds for i = eachindex(x,y,z)
        z[i] = x[i] + y[i]
    end
    df.z = z
    nothing
end
function broadcast!(df)
    df.z = df.x .+ df.y
    nothing
end
@btime loop!($df)
@btime broadcast!($df)

```

I get

```julia
julia> @btime loop!($df)
  104.254 μs (5124 allocations: 104.16 KiB)

julia> @btime broadcast!($df)
  1.188 μs (4 allocations: 8.20 KiB)

```

My point is that the broadcast itself is a function barrier.  
A loop is not a function barrier.

Thus, if the broadcast is type stable, it will be fast even if the inputs are not. There will be a slow dynamic dispatch to the fast `::Vector{Float64} .+ ::Vector{Float64}`.  
However, a loop doesn’t have this barrier, so it suffers from dynamic dispatches on every iteration.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 17, 2023, 7:00pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/15 "2023-02-17T19:00:03Z")

</div>

> [@zeroexcuses](#):
>
> I’m new to Julia. Can you please explain when vectorizing would be SLOWER than looping? Intuition says vectorizing (when possible) is always faster.

“vectorizing” in terms of broadcasting will often fail to “vectorize” in the low level sense even when loops would, because broadcasting is fundamentally dynamic.  
This is why `FastBroadcast.jl` has the option `broadcast=false`, which defaults to `false`, meaning it won’t actually broadcast.  
This gives access to the nice syntax of broadcasting and the ability to still support GPUs, but without the occasionally-performance-killing dynamism of `broadcast=true`.  
Even if you set `broadcast=true`, `FastBroadcast.jl` will still special case the common situation of not broadcasting, and generate specialized code (loops) for only that case, falling back to a slower implementation if necessary.

Anyway, my point is that vectorized syntax by default is bad for vectorization. You can use `FastBroadcast.jl` to remedy this, which the SciML ecosystem does.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [February 17, 2023, 7:01pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/16 "2023-02-17T19:01:25Z")

</div>

Just a meta comment: the use of the word “vectorization” in programming is pretty overloaded. To some people, it means something along the lines of “functions acting on an entire array” and to others it means [SIMD](https://en.wikipedia.org/wiki/Single_instruction,_multiple_data).

In julia circles, people typically mean SIMD when they say ‘vectorization’ but this is different from say Python, Matlab, and R where typically the only way to access SIMD is through their ‘vectorized’ functions, whereas in julia it’s pretty easy to get a loop to use SIMD instructions (it’ll often just happen automatically).

This can sometimes cause people to get confused on both sides, so something to be aware of.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 17, 2023, 7:33pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/17 "2023-02-17T19:33:38Z")

</div>

@Elrod, thanks for the very educational example using DataFrames, which are type-unstable, making your point very clear. Very appreciated.

As a side note, if we create a function barrier with the loop then it would be even faster than the vectorized version:

```julia
function sumxy!(z, x, y)
    @inbounds for i = eachindex(x,y,z)
        z[i] = x[i] + y[i]
    end
    return z
end
function loop!(df)
    (; x, y) = df
    z = similar(x)
    df.z = sumxy!(z, x, y) # function barrier
    nothing
end

```

---

<div class="post-metadata">

**Author:** ![Eben60](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eben60/32/13475_2.png) [@Eben60](https://discourse.julialang.org/u/Eben60)\
**Post date:** [February 17, 2023, 8:26pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/18 "2023-02-17T20:26:14Z")

</div>

> [@rafael.guerra](#):
>
> @Elrod, thanks for the very educational example using DataFrames, which are type-unstable, making your point very clear.

Can anybody explain why the type instability? I’d naively think, `df` is passed as a function argument, `eltype` of the `df` columns is `Float64`

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 17, 2023, 9:29pm UTC](https://discourse.julialang.org/t/loops-vs-vectorization-when-to-use/94724/19 "2023-02-17T21:29:27Z")

</div>

@Eben60, see @bkamins multiple posts and courses on this matter, including [here](https://github.com/bkamins/Julia-DataFrames-Tutorial/blob/master/11_performance.ipynb), and his [latest book](https://discourse.julialang.org/t/julia-for-data-analysis-book/78090/11) that covers this issue nicely.

Basically, the columns of a DataFrame are AbstractVectors (for maximal user flexibility) and the compiler does not know the actual memory layout. However, inside the function barrier, everything becomes type stable for the compiler and faster code can be generated.
