# CartesianIndices performance: 3D faster than 2D?

**URL:** <https://discourse.julialang.org/t/cartesianindices-performance-3d-faster-than-2d/94110>\
**Category:** Performance\
**Tags:** cartesianindices\
**Created:** [February 6, 2023, 12:06am UTC](https://discourse.julialang.org/t/cartesianindices-performance-3d-faster-than-2d/94110 "2023-02-06T00:06:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![lntricate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lntricate/32/46878_2.png) [@lntricate](https://discourse.julialang.org/u/lntricate)\
**Post date:** [February 6, 2023, 12:06am UTC](https://discourse.julialang.org/t/cartesianindices-performance-3d-faster-than-2d/94110/1 "2023-02-06T00:06:05Z")

</div>

I have encountered a strange occurence when testing the performance of `CartesianIndices`:

```julia
using BenchmarkTools

function test(t::Tuple)
  for _ ∈ CartesianIndices(t)
  end
end

@btime test((100,)) # 3.094 ns
@btime test((100, 100)) # 9.531 μs
@btime test((100, 100, 100)) # 7.267 μs
@btime test((100, 100, 100, 100)) # 152.958 ms

```

There are 2 behaviors here I’m confused about:

1. The 3-dimensional loop takes less time than the 2-dimensional loop.
2. The 4-dimensional loop takes over 20,000 times longer than the 3-dimensional loop despite (theoretically) only having to iterate over 100x the indices. Why does it not take 100x longer?

Any help would be appreciated!

---

<div class="post-metadata">

**Author:** ![Paulo\_Jabardo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paulo_jabardo/32/3196_2.png) [@Paulo\_Jabardo](https://discourse.julialang.org/u/Paulo_Jabardo)\
**Post date:** [February 6, 2023, 4:21pm UTC](https://discourse.julialang.org/t/cartesianindices-performance-3d-faster-than-2d/94110/2 "2023-02-06T16:21:45Z")

</div>

I don’t know what the compiler is doing but remember that your code is not doing anything and it _might_ be optimized away.

For the first two cases, the for body is optimized away:

```julia-repl
julia> @code_llvm test((100,))
; @ REPL[2]:1 within `test`
define void @julia_test_568([1 x i64]* nocapture nonnull readonly align 8 dereferenceable(8) %0) #0 {
top:
; @ REPL[2]:3 within `test`
  ret void
}

```

```julia-repl
julia> @code_llvm test((100,100))
; @ REPL[2]:1 within `test`
define void @julia_test_574([2 x i64]* nocapture nonnull readonly align 8 dereferenceable(16) %0) #0 {
top:
; @ REPL[2]:3 within `test`
  ret void
}

```

Now, for `test(100,100,100)`, the code is much larger.

By the way, my timings are:

```julia-repl
julia> @btime test((100,))
  2.194 ns (0 allocations: 0 bytes)

julia> @btime test((100,100))
  2.194 ns (0 allocations: 0 bytes)

julia> @btime test((100,100,100))
  3.444 μs (0 allocations: 0 bytes)

julia> @btime test((100,100,100,100))
  60.121 ms (0 allocations: 0 bytes)

```

Paulo

---

<div class="post-metadata">

**Author:** ![lntricate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lntricate/32/46878_2.png) [@lntricate](https://discourse.julialang.org/u/lntricate)\
**Post date:** [February 6, 2023, 8:16pm UTC](https://discourse.julialang.org/t/cartesianindices-performance-3d-faster-than-2d/94110/3 "2023-02-06T20:16:52Z")

</div>

Thanks for the response! Indeed the compiler optimizing away the loop body explains all of these results. After testing the same function with `@noinline` and a simple addition in the loop body, I got the following:

```julia
@btime test((100,)) # 112.413 ns
@btime test((100, 100)) # 12.523 μs
@btime test((100, 100, 100)) # 1.094 ms
@btime test((100, 100, 100, 100)) # 139.926 ms

```

Each roughly 100 times more than the previous one as expected.
