# Why are for loops slower than broadcast with splatted arguments?

**URL:** <https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801>\
**Category:** New to Julia\
**Tags:** performance, tuple\
**Created:** [November 6, 2019, 6:02pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801 "2019-11-06T18:02:42Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![jlchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlchan/32/10958_2.png) [@jlchan](https://discourse.julialang.org/u/jlchan)\
**Post date:** [November 6, 2019, 6:02pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/1 "2019-11-06T18:02:43Z")

</div>

I’m broadcasting functions which return tuples, and have been converting arrays of tuples to tuples of arrays using “unzip”. I used for loops over each individual entry and re-assembled the output as a tuple.

```julia
using BenchmarkTools

function foo(a,b,c,d)
    ab = exp(a*b)
    return ab+c, ab-d
end

N = 10000
a,b,c,d = [randn(N) for i = 1:4]
X = (a,b)
Y = (c,d)

function foo_loop(N,a,b,c,d)
    for i = 1:N
        foo(a[i],b[i],c[i],d[i])
    end
end
@btime foo_loop($N,$a,$b,$c,$d)

function foo_loop_splat!(N,X,Y,X_entry,Y_entry)
    for i = 1:N
        for fld in eachindex(X)
            X_entry[fld] = X[fld][i]
            Y_entry[fld] = Y[fld][i]
        end
        foo(X_entry...,Y_entry...)
    end
end
X_entry, Y_entry = [zeros(length(X)) for i = 1:2]
@btime foo_loop_splat!($N,$X,$Y,$X_entry,$Y_entry)

```

I noticed that the splatted for loop is much slower than the broadcasted “unzip” approach

```julia
  102.436 μs (0 allocations: 0 bytes)
  1.046 ms (50000 allocations: 937.50 KiB)

```

In comparison, using broadcasting and unzip

```julia
unzip(a) = map(x->getfield.(a, x), fieldnames(eltype(a)))
@btime unzip(foo.($X...,$Y...))
@btime unzip(foo.($a,$b,$c,$d))

```

both runtimes are about the same

```julia
  158.962 μs (8 allocations: 312.78 KiB)
  156.331 μs (8 allocations: 312.78 KiB)

```

Why does this occur? I’m not clear on why splatting adds so much extra time. Is it [runtime dispatch](https://discourse.julialang.org/t/splatting-arguments-causes-30x-slow-down/16964)?

---

<div class="post-metadata">

**Author:** ![rdeits](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rdeits/32/286_2.png) [@rdeits](https://discourse.julialang.org/u/rdeits)\
**Post date:** [November 6, 2019, 6:10pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/2 "2019-11-06T18:10:29Z")

</div>

You’re benchmarking in global scope, so the performance numbers are not particularly meaningful. See [Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/index.html#Avoid-global-variables-1) and [GitHub - JuliaCI/BenchmarkTools.jl: A benchmarking framework for the Julia language](https://github.com/JuliaCI/BenchmarkTools.jl#quick-start)

---

<div class="post-metadata">

**Author:** ![Pbellive](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pbellive/32/3604_2.png) [@Pbellive](https://discourse.julialang.org/u/Pbellive)\
**Post date:** [November 6, 2019, 6:15pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/3 "2019-11-06T18:15:30Z")

</div>

Just a quck addition to @rdeits response. The benchmarking in global scope really slows things down the for loop code here. On my machine, I get:

```julia
using BenchmarkTools

function foo(a,b,c,d)
    ab = exp(a*b)
    return ab+c, ab-d
end

N = 1000
a,b,c,d = [randn(N) for i = 1:4]
X = (a,b)
Y = (c,d)

u,v = [zeros(N) for i = 1:2]
@btime begin
    for i = 1:N
        foo(a[i],b[i],c[i],d[i])        
    end

```

Then, if I define

```julia
function foo_loop(N,a,b,c,d)
         for i = 1:N
               foo(a[i],b[i],c[i],d[i])        
           end
end

109.880 μs

```

Then benchmark by calling the function and interpolating the input arguments into the benchmarking expression ([https://github.com/JuliaCI/BenchmarkTools.jl/blob/master/doc/manual.md#interpolating-values-into-benchmark-expressions](https://github.com/JuliaCI/BenchmarkTools.jl/blob/master/doc/manual.md#interpolating-values-into-benchmark-expressions)) I get:

```julia
julia> @btime foo_loop($N,$a,$b,$c,$d)
  7.273 μs (0 allocations: 0 bytes)

```

---

<div class="post-metadata">

**Author:** ![jlchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlchan/32/10958_2.png) [@jlchan](https://discourse.julialang.org/u/jlchan)\
**Post date:** [November 6, 2019, 6:23pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/4 "2019-11-06T18:23:52Z")

</div>

Thanks @rdeits and @Pbellive! I’ve edited the original code snippet to interpolate values and avoid global scope timing.

```julia
function foo_loop(N,a,b,c,d)
    for i = 1:N
        foo(a[i],b[i],c[i],d[i])
    end
end
@btime foo_loop($N,$a,$b,$c,$d)

function foo_loop_splat!(N,X,Y,X_entry,Y_entry)
    for i = 1:N
        for fld in eachindex(X)
            X_entry[fld] = X[fld][i]
            Y_entry[fld] = Y[fld][i]
        end
        foo(X_entry...,Y_entry...)
    end
end
X_entry, Y_entry = [zeros(length(X)) for i = 1:2]
@btime foo_loop_splat!($N,$X,$Y,$X_entry,$Y_entry)

unzip(a) = map(x->getfield.(a, x), fieldnames(eltype(a)))
@btime unzip(foo.($X...,$Y...))
@btime unzip(foo.($a,$b,$c,$d))

```

The for loop is now faster than the unzip-and-broadcast code, though the splatting is still slower.

```julia
  158.962 μs (8 allocations: 312.78 KiB)
  156.331 μs (8 allocations: 312.78 KiB)
  102.436 μs (0 allocations: 0 bytes)
  1.046 ms (50000 allocations: 937.50 KiB)

```

---

<div class="post-metadata">

**Author:** ![jlchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlchan/32/10958_2.png) [@jlchan](https://discourse.julialang.org/u/jlchan)\
**Post date:** [November 6, 2019, 7:55pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/5 "2019-11-06T19:55:43Z")

</div>

I think I figured out why it’s so slow - I wasn’t being careful with tuples and arrays.

In the original splatted loop code, I’m splatting the X\_entry, Y\_entry arguments, which (I guess) converts from an Array to a Tuple.

```julia
function foo_loop_splat!(N,X,Y,Xi,Yi)
    for i = 1:N
        for fld in eachindex(X)
            Xi[fld] = X[fld][i]
            Yi[fld] = Y[fld][i]
        end
        foo(Xi...,Yi...)
    end
end

```

An old post by @rdeits on [Array of tuples - #7 by rdeits](https://discourse.julialang.org/t/array-of-tuples/11435/7) noted that tuples have much less overhead compared to arrays. If I replace the for loop with the (simpler) tuple-based version

```julia
function foo_loop_splat!(N,X,Y)
    for i = 1:N
        foo(getindex.(X,i)...,getindex.(Y,i)...)
    end
end

```

then the non-splatted and splatted for loop give identical runtimes

```julia
@btime foo_loop($N,$a,$b,$c,$d)
@btime foo_loop_splat!($N,$X,$Y)

```

```julia
  1.049 ms (0 allocations: 0 bytes)
  1.093 ms (0 allocations: 0 bytes)

```

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [November 6, 2019, 7:59pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/6 "2019-11-06T19:59:34Z")

</div>

The key is that arrays don’t encode their lengths into the type system like tuples do. This means when you splat it into a function call, Julia doesn’t know (ahead of time) how many arguments there will be — and thus cannot predict which method will be called. This is essentially a type instability.

---

<div class="post-metadata">

**Author:** ![jlchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlchan/32/10958_2.png) [@jlchan](https://discourse.julialang.org/u/jlchan)\
**Post date:** [November 12, 2019, 9:30pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/7 "2019-11-12T21:30:26Z")

</div>

@mbauman - thanks for the interpretation as type instability! That’s helpful to know.

I converted my code to use tuples of arrays instead of arrays of arrays. However, the immutability of tuples prevents me from modifying them. Is there a recommended mutable alternative to tuples (StructArrays or StaticArrays?) which also encode lengths into their type?

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [November 12, 2019, 9:36pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/8 "2019-11-12T21:36:20Z")

</div>

You could try using an `MArray` from StaticArrays.jl: [API · StaticArrays.jl](https://juliaarrays.github.io/StaticArrays.jl/latest/pages/api/#StaticArrays.MArray) since they are mutable and fixed size.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [November 12, 2019, 9:38pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/9 "2019-11-12T21:38:40Z")

</div>

I don’t think splatting `MArray`s will be as fast as tuples, though. The type-system-encoded length is necessary but not sufficient. I think tuples have some special builtin secret sauce — or they did last I checked.

It’s often possible to simply avoid splatting (or not worry about it).

---

<div class="post-metadata">

**Author:** ![Daniel\_Berge](https://avatars.discourse-cdn.com/v4/letter/d/eb9ed0/32.png) [@Daniel\_Berge](https://discourse.julialang.org/u/Daniel_Berge)\
**Post date:** [November 12, 2019, 9:45pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/10 "2019-11-12T21:45:45Z")

</div>

I’m curious why this is. A `StaticArray` is just a `NTuple` underneath. Is it just a missing optimization?

---

<div class="post-metadata">

**Author:** ![jlchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlchan/32/10958_2.png) [@jlchan](https://discourse.julialang.org/u/jlchan)\
**Post date:** [November 12, 2019, 9:54pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/11 "2019-11-12T21:54:04Z")

</div>

Splatting tuples does seem much faster than splatting mutable or sized arrays. Running the same tests with SizedVectors

```julia
# replace tuples with sized vectors
X = SizedVector{2}(a,b)
Y = SizedVector{2}(c,d)
@btime foo_loop_splat!($N,$X,$Y)

```

is pretty slow compared to splatting tuples.

```julia
  106.395 μs (0 allocations: 0 bytes) # splatted tuples
  4.501 ms (170000 allocations: 5.19 MiB) # splatted SizedVector

```

I’ve gotten around immutability by overwriting tuples in global scope, e.g.

```julia
X = X .+ Y

```

It doesn’t seem that slow to me, but I can see where it would be limiting.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [November 12, 2019, 9:54pm UTC](https://discourse.julialang.org/t/why-are-for-loops-slower-than-broadcast-with-splatted-arguments/30801/12 "2019-11-12T21:54:44Z")

</div>

Splatting is “just” iteration, though, and that’s a tough thing to “see through” to the end in general at the type system level.
