# Performance drop for explicit indexing in fused loops

**URL:** <https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459>\
**Category:** General Usage\
**Created:** [June 26, 2017, 5:23am UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459 "2017-06-26T05:23:27Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![fjarri](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fjarri/32/630_2.png) [@fjarri](https://discourse.julialang.org/u/fjarri)\
**Post date:** [June 26, 2017, 5:23am UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/1 "2017-06-26T05:23:27Z")

</div>

Consider the following code doing a simple computation:

```julia
function test_vectorized(a, a_old)
    @. a[:] = a_old[:] * 2
end

function test_vectorized2(a, a_old)
    @. a = a_old * 2
end

function test_devectorized(a, a_old)
    nx = size(a, 1)
    @inbounds @simd for i = 1:nx
        a[i] = a_old[i] * 2
    end
end

a = rand(100000000)

a1 = copy(a)
test_vectorized(a1, a)
@time test_vectorized(a1, a)

a2 = copy(a)
test_vectorized2(a2, a)
@time test_vectorized2(a2, a)

a3 = copy(a)
test_devectorized(a3, a)
@time test_devectorized(a3, a)

@assert isapprox(a1, a2)
@assert isapprox(a1, a3)

```

In Julia 0.6 the first vectorized function is about 5 times slower:

```julia
> julia test.jl
  0.617197 seconds (86 allocations: 762.946 MiB, 17.37% gc time)
  0.127845 seconds (4 allocations: 160 bytes)
  0.125778 seconds (4 allocations: 160 bytes)

```

It would seem that there is enough information in `test_vectorized()` for the compiler to turn it into something similar to `test_devectorized()`, but, apparently, it is not the case. `test_vectorized2()`, on the other hand, shows the same performance. I guess the explicit indexing in `test_vectorized()` is the reason. Is there some way to improve the performance of fused loops in case of explicit indexing? In my actual program I am performing a stencil-type calculation, so I need a way to specify ranges (i.e., I have something like

```julia
@. a[2:end-1, 2:end-1] = (
    (a_old[1:end-2, 2:end-1] + a_old[3:end, 2:end-1]) * dyd
    + (a_old[2:end-1,1:end-2] + a_old[2:end-1, 3:end]) * dxd)

```

)

---

<div class="post-metadata">

**Author:** ![aaowens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaowens/32/12101_2.png) [@aaowens](https://discourse.julialang.org/u/aaowens)\
**Post date:** [June 26, 2017, 5:45am UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/2 "2017-06-26T05:45:00Z")

</div>

```julia
@views function test_vectorized3(a, a_old)
    @. a[:] = a_old[:] * 2
end
julia> @time test_vectorized3(a2, a);
  0.126813 seconds (5 allocations: 208 bytes)

```

I think views will also work for your stencils.

---

<div class="post-metadata">

**Author:** ![fjarri](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fjarri/32/630_2.png) [@fjarri](https://discourse.julialang.org/u/fjarri)\
**Post date:** [June 26, 2017, 6:51am UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/3 "2017-06-26T06:51:17Z")

</div>

Thanks, that did help a lot! The performance is the same for the simple operation from the opening post, and the performance of a stencil is close to the devectorized version, although still not exactly the same (~25% slower):

```julia
function test_vectorized(a, a_old)
    @views @. a[2:end-1, 2:end-1] = (
        (a_old[1:end-2, 2:end-1] + a_old[3:end, 2:end-1]) * 0.1
        + (a_old[2:end-1,1:end-2] + a_old[2:end-1, 3:end]) * 0.2)
end

function test_devectorized(a, a_old)
    nx, ny = size(a)
    @inbounds @simd for j = 2:(ny-1)
        for i = 2:(nx-1)
            a[i,j] = ((a_old[i-1, j] + a_old[i+1, j]) * 0.1 
                + (a_old[i, j-1] + a_old[i, j+1]) * 0.2)
        end
    end
end

a = rand(10000, 10000)

a1 = copy(a)
test_vectorized(a1, a)
@time test_vectorized(a1, a)

a2 = copy(a)
test_devectorized(a2, a)
@time test_devectorized(a2, a)

@assert isapprox(a1, a2)

```

This gives

```julia
> julia test.jl
  0.218883 seconds (87 allocations: 6.623 KiB)
  0.152084 seconds (4 allocations: 160 bytes)

```

(without the `@views`, `test_vectorized()` is about 5 times slower).

---

<div class="post-metadata">

**Author:** ![traktofon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/traktofon/32/591_2.png) [@traktofon](https://discourse.julialang.org/u/traktofon)\
**Post date:** [June 26, 2017, 9:22am UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/4 "2017-06-26T09:22:55Z")

</div>

> [@fjarri](#):
>
> \> julia test.jl  
> 0.218883 seconds (87 allocations: 6.623 KiB)  
> 0.152084 seconds (4 allocations: 160 bytes)

It seems that your first `@time` call is also timing the compilation of `@time` itself (or more precisely, of what it expands to), hence the much larger number of allocations. In general, it is recommended to use the package `BenchmarkTools.jl` for benchmarking, as follows:

```julia
julia> using BenchmarkTools

julia> @btime test_vectorized($a1, $a);
  215.568 ms (4 allocations: 256 bytes)

julia> @btime test_devectorized($a2, $a);
  187.499 ms (0 allocations: 0 bytes)

```

So there is indeed some overhead to the vectorized/views version, which I think currently cannot be avoided, because each view needs an allocation (although a small one).

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [June 26, 2017, 1:38pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/5 "2017-06-26T13:38:21Z")

</div>

Don’t benchmark in global scope.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [June 26, 2017, 1:41pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/6 "2017-06-26T13:41:34Z")

</div>

On Julia 0.6, you can also use the new syntax `a .= a_old .* 2`, which should perform quite well as a replacement for `a[:] = a_old[:] * 2`.

---

<div class="post-metadata">

**Author:** ![fjarri](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fjarri/32/630_2.png) [@fjarri](https://discourse.julialang.org/u/fjarri)\
**Post date:** [June 26, 2017, 1:52pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/7 "2017-06-26T13:52:09Z")

</div>

@dpsanders, could you elaborate on that a bit?

@nalimilan, note the `@.` in front of the expression, it converts operators into broadcasted form automatically. The question was about the difference between `a` and `a[:]`.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [June 26, 2017, 2:09pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/8 "2017-06-26T14:09:44Z")

</div>

> [@fjarri](#):
>
> @nalimilan, note the @. in front of the expression, it converts operators into broadcasted form automatically. The question was about the difference between a and a[:].

I know that, but in Julia 0.6 you can directly use `.=`, there’s no reason to use `@.` or `a[:]` for this kind of thing anymore.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [June 26, 2017, 2:17pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/9 "2017-06-26T14:17:45Z")

</div>

`@views @.` is fully inplace. Asymtopically it’ll do as well as the devectorized form. Currently it has overhead because views are not stack-allocated, but that’s just a missing compiler optimization.

> [@nalimilan](#):
>
> but in Julia 0.6 you can directly use .=, there’s no reason to use @. or a[:] for this kind of thing anymore.

I agree there’s no reason to `a[:]` anymore, but `a .= @. ...` seems weird and is no more efficient than `@. a = ...`

---

<div class="post-metadata">

**Author:** ![fjarri](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fjarri/32/630_2.png) [@fjarri](https://discourse.julialang.org/u/fjarri)\
**Post date:** [June 26, 2017, 2:24pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/10 "2017-06-26T14:24:44Z")

</div>

`@.` is useful if you have a lot of dotted operators in the expression (I have only two, so it’s not that evident, but this example was reduced from a bigger one; see my post with the stencil code).

`[:]` is crucial to the question; I was worried about the performance of slices. Of course, `a` and `a[:]` give the same result (for a 1D array), but even a trivial slice `[:]` still reproduces the issue with performance. It’s a MWE, I tried to make it as simple as possible. In my actual code I have non-trivial slices, of course.

---

<div class="post-metadata">

**Author:** ![fjarri](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fjarri/32/630_2.png) [@fjarri](https://discourse.julialang.org/u/fjarri)\
**Post date:** [June 26, 2017, 2:25pm UTC](https://discourse.julialang.org/t/performance-drop-for-explicit-indexing-in-fused-loops/4459/11 "2017-06-26T14:25:52Z")

</div>

> [@ChrisRackauckas](#):
>
> @views @. is fully inplace. Asymtopically it’ll do as well as the devectorized form. Currently it has overhead because views are not stack-allocated, but that’s just a missing compiler optimization.

Glad to hear it, I was worried that I missed some kind of compiler hint. Although I do not really see how several small allocations can take so much time. I already have a 10^10-element array, when does it start to even out?

> [@ChrisRackauckas](#):
>
> but a .= @. … seems weird and is no more efficient than @.

I honestly do not see this construction used anywhere in this thread. I only have `@. a = ...`.
