# Question: Improve the speed of for loop

**URL:** <https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204>\
**Category:** Performance\
**Tags:** question\
**Created:** [April 7, 2023, 7:08am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204 "2023-04-07T07:08:33Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 7, 2023, 7:08am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/1 "2023-04-07T07:08:33Z")

</div>

I have written massive ‘for loop’ code as follows:

```julia
Meta_start=390; Meta_end=400; Meta_div=trunc(Int,((Meta_end-Meta_start)*100)); #temporary
        Meta_cal=LinRange(Meta_start,Meta_end,Meta_div); k=1:1:Meta_div #::Vector{Int64}
            I_sharp=Array{Float64,2}(undef,150,27);
            I_broad=Array{Float64,1}(undef,Meta_div);

            Re=rand(5,11); FCF=rand(5,11); IB=0.1;
            Meta=rand(150,27,55); Meta_shift=zeros(55);

            for v_C=1:1:5
                for v_B=1:1:11
                    numb_v=(v_B+11*(v_C-1))
                    for bra=1:1:27
                         for J1=1:1:150 # 150
                            I_sharp[J1,bra]=Re[v_C,v_B]^2+*(1/(Meta[J1,bra,numb_v].-Meta_shift[numb_v])^4)*FCF[v_C,v_B]
                            I_broad[k]+=I_sharp[J1,bra].*exp.(-2*(Meta_cal[k].-(Meta[J1,bra,numb_v].-Meta_shift[numb_v])).^2/IB)
                        end 
                    end
                end
            end

```

It took 3.448 seconds with 8691813 allocations and 10.30 GiB on my computer.  
I tried to use @inbounds @simd and mul!(C,A,B), but it didn’t work, or I guess I used it the wrong way.  
How can I improve the speed of this code?

---

<div class="post-metadata">

**Author:** ![Tarny\_GG\_Channie](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Tarny\_GG\_Channie](https://discourse.julialang.org/u/Tarny_GG_Channie)\
**Post date:** [April 7, 2023, 11:47am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/2 "2023-04-07T11:47:42Z")

</div>

Some parts can probably be vectorized using [loop vectorization](https://github.com/JuliaSIMD/LoopVectorization.jl), the `a.+b` or similar vectorizations are convenient but not efficient, do not vectorize the code this way if you care about performance. Also, see  
[Transducers.jl](https://juliafolds.github.io/Transducers.jl/dev/) or other packages from the Juliafolds to efficiently do the reduction of I\_broad.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 7, 2023, 12:04pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/3 "2023-04-07T12:04:06Z")

</div>

are you benchmarking in global scope? put this in a function.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 7, 2023, 2:31pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/4 "2023-04-07T14:31:12Z")

</div>

> [@manu\_han](#):
>
> ```julia
> for v_C=1:1:5
> for v_B=1:1:11
> 
> ```

Hardcoding iteration limits isn’t a great idea, you should rather use `axes` or `eachindex`. Anyway, use `1:5`, not `1:1:5`, otherwise the compiler cannot be certain about the stepsize in the type domain.

But most importantly, put your code in a function, and avoid global variables. And definitely read the [performance tips](https://docs.julialang.org/en/v1/manual/performance-tips/)

---

<div class="post-metadata">

**Author:** ![PeterSimon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petersimon/32/25193_2.png) [@PeterSimon](https://discourse.julialang.org/u/PeterSimon)\
**Post date:** [April 7, 2023, 2:49pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/5 "2023-04-07T14:49:24Z")

</div>

You have `+*` in one of your expressions. Not sure what you intended here, but I believe the asterisk has no effect since it is followed by a parentheses containing a single quantity, interpreted as a product of a set of one numbers, ie, the number itself.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 7, 2023, 2:54pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/6 "2023-04-07T14:54:20Z")

</div>

> [@manu\_han](#):
>
> `I_broad[k]+=I_sharp[J1,bra].*exp.(-2*(Meta_cal[k].-(Meta[J1,bra,numb_v].-Meta_shift[numb_v])).^2/IB)`

Aside: As I understand it, all of these terms are scalars, so the dots do nothing. (They mostly don’t hurt, but you might as well use `*, - , ^` instead of `.*, .-, .^`.)

---

<div class="post-metadata">

**Author:** ![PeterSimon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petersimon/32/25193_2.png) [@PeterSimon](https://discourse.julialang.org/u/PeterSimon)\
**Post date:** [April 7, 2023, 7:40pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/7 "2023-04-07T19:40:48Z")

</div>

This code has several problems but one is extremely serious: `I_broad` is never initialized. It is incremented in the innermost loop but since it is never initialized, it contains random garbage. This should be fixed before you ask for any more help with optimizing its speed.

---

<div class="post-metadata">

**Author:** ![Tarny\_GG\_Channie](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Tarny\_GG\_Channie](https://discourse.julialang.org/u/Tarny_GG_Channie)\
**Post date:** [April 7, 2023, 8:58pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/8 "2023-04-07T20:58:03Z")

</div>

I think it’s okay actually.

```julia
arr = rand(1000000)
function test_range_1(arr)
    ans = 0
    for i in 1:1:1000000
        ans += arr[i]
    end
    ans
end
function test_range_2(arr)
    ans = 0
    for i in 1:1000000
        ans += arr[i]
    end
    ans
end

using BenchmarkTools

@btime test_range_1(arr)
@btime test_range_2(arr)

```

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 8, 2023, 6:43am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/10 "2023-04-08T06:43:47Z")

</div>

Originally, this function was inside the function, and the benchmark time I posted was run inside the function.

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 8, 2023, 6:46am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/11 "2023-04-08T06:46:02Z")

</div>

Yeah, I’m reading the performance tips. and unfortunately, using 1:5 was not make my code faster :(.

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 8, 2023, 7:52am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/12 "2023-04-08T07:52:27Z")

</div>

I minimized the vectorization following your comment, but the effect is little.  
and Thank you for introducing the interesting package.  
It seems that [`Transducers.Scan`] and [`Transducers.Iterated`] could be applied in my code.  
could you give me an example? I have no idea how to use it on I\_broad.

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 8, 2023, 7:53am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/13 "2023-04-08T07:53:51Z")

</div>

> [@PeterSimon](#):
>
> You have `+*` in one of your expressions. Not sure what you intended here, but I believe the asterisk has no effect since it is followed by a parentheses containing a single quantity, interpreted as a product of a set of one numbers, ie, the number itself.

Thanks for pointing out the typo. The original intent was \*.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 8, 2023, 10:31am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/15 "2023-04-08T10:31:44Z")

</div>

> [@manu\_han](#):
>
> unfortunately, using 1:5 was not make my code faster :(.

It’s more about style than performance, though. `1:5` is nicer. But, really, you should use `eachindex` or `axes`. Hardcoding iteration limits is very error-prone.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [April 8, 2023, 10:32am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/16 "2023-04-08T10:32:18Z")

</div>

```julia
I_broad=zeros(Float64, Meta_div) # zeros, not random memory
...
# fixed .* typo, got rid of dotted operator for scalars (no difference)
I_sharp[J1,bra]=Re[v_C,v_B]^2*(1/(Meta[J1,bra,numb_v]-Meta_shift[numb_v])^4)*FCF[v_C,v_B]

# @view prevents .+= from allocating I_broad[k] on the right side
# dot all the functions over arrays to fuse into 1 in-place loop
@view(I_broad[k]) .+= I_sharp[J1,bra].*exp.(-2 .*( Meta_cal[k] .-(Meta[J1,bra,numb_v]-Meta_shift[numb_v])).^2 ./IB)

```

These reduce the allocations to 8 and 1.74MiB in a function on v1.8.5; I’m not sure why for 2 of them, only 6 things are allocated before the loop (EDIT it seems to be an effect of filling arrays. If you omit the loop and change all `undef` to `zeros`, you get those allocation numbers. If you omit the loop and use the original 2 `undef` lines, you get 5 allocations 1.701 MiB.) I just noticed that the `k` was not a scalar index so I had to pay attention to where any `Array`s were being allocated e.g. the `LinRange` at unfused `.^`. Not sure why you’re slicing with `k=1:1:Meta_div` when the arrays seem to be length `Meta_div` already.

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 9, 2023, 7:57am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/17 "2023-04-09T07:57:21Z")

</div>

The code performance improve from 1.97 sec (1336598 allocations: 10.12 GiB) to 1.32 sec (8 allocations: 1.74 MiB)!  
The reason for doing ‘slicing k’ was for the vectorization of I\_broad, and it seems unnecessary.

```julia
I_broad+= I_sharp[J1,bra].*exp.(-2 .*( Meta_cal .-(Meta[J1,bra,numb_v]-Meta_shift[numb_v])).^2 ./IB)

```

It reduced the time but still requires a lot of memory and allocation.  
In this case, I think I need to use something other than the @view.

Thank you for the detailed example, Benny. it is really helpful.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [April 9, 2023, 8:24am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/18 "2023-04-09T08:24:15Z")

</div>

The reason `@view` is used for slices is because `I_broad[k]` on the right side of `I_broad[k] .= I_broad[k] .+ ...` is allocating a copy. You don’t need `@view` if you don’t slice because `I_broad` on the right side of `I_broad .= I_broad .+ ...` reuses the array, no allocation needed. Try this line and see if you reduce the allocations again:

```julia
I_broad .+= I_sharp[J1,bra].*exp.(-2 .*( Meta_cal .-(Meta[J1,bra,numb_v]-Meta_shift[numb_v])).^2 ./IB)

```

> [@manu\_han](#):
>
> 1.97 sec … to 1.32 sec

I’m getting the feeling that your code does so much number crunching that it vastly outweighs the time spent allocating and deallocating, so reducing allocations has a small improvement on runtime.

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 10, 2023, 1:46am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/19 "2023-04-10T01:46:41Z")

</div>

I have tried the code you suggested, and it shows a 20 % speed-up compared to mine.  
Even though the number of allocations did not decrease, memory did decrease 12 times.  
A period ‘.’ made it like this. (reducing process time and memory)  
It seems trivial, but important… I’ll have to pay attention it for optimization.  
Thank you. 🙂

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [April 10, 2023, 1:56am UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/20 "2023-04-10T01:56:13Z")

</div>

> [@manu\_han](#):
>
> the number of allocations did not decrease

Odd, the `.+=` compared to `+=` would save an allocation of an array on the right side. I quickly ran the code with the `I_broad+=` line versus the `I_broad .+=` line, it went from 445510 to the 8 allocations (1.74 MiB) as expected.

> [@manu\_han](#):
>
> memory did decrease 12 times.

I cannot explain the memory decreasing by 12 times, that memory is all allocated before the loop, so did you edit the arrays to be smaller?

---

<div class="post-metadata">

**Author:** ![manu\_han](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/manu_han/32/47520_2.png) [@manu\_han](https://discourse.julialang.org/u/manu_han)\
**Post date:** [April 10, 2023, 12:19pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/22 "2023-04-10T12:19:55Z")

</div>

Maybe it is due to the code read the data using readdlm() not using rand(), in real code.  
But I don’t know why the two cases are not the same, both of them is float 64 type and the same size array. (readdlm demand more allocations than rand()).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [April 10, 2023, 5:08pm UTC](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204/23 "2023-04-10T17:08:25Z")

</div>

```julia
julia> @benchmark f1!($I_sharp, $I_broad, $Re, $Meta, $Meta_shift, $FCF, $Meta_cal, $k, $IB)
BenchmarkTools.Trial: 1 sample with 1 evaluation.
 Single result which took 2.981 s (0.00% GC) to evaluate, with a memory estimate of 0 bytes, over 0 allocations.

julia> @benchmark f2!($I_sharp, $I_broad, $Re, $Meta, $Meta_shift, $FCF, $Meta_cal, $k, $IB)BenchmarkTools.Trial: 9 samples with 1 evaluation.
 Range (min … max): 120.850 ms … 121.046 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 120.854 ms ┊ GC (median): 0.00% Time (mean ± σ): 120.877 ms ± 63.380 μs ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▃█                                                             
  ██▇▇▇▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▇ ▁
  121 ms Histogram: frequency by time 121 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

`f2!` uses `@turbo`.

```julia
Meta_start=390; Meta_end=400; Meta_div=trunc(Int,((Meta_end-Meta_start)*100)); #temporary
Meta_cal=LinRange(Meta_start,Meta_end,Meta_div); k=1:Meta_div #::Vector{Int64}
I_sharp=Array{Float64,2}(undef,150,27);
I_broad=Array{Float64,1}(undef,Meta_div);

Re=rand(5,11); FCF=rand(5,11); IB=0.1;
Meta=rand(150,27,55); Meta_shift=zeros(55);

function f1!(I_sharp, I_broad, Re, Meta, Meta_shift, FCF, Meta_cal, K, IB)
  @inbounds @fastmath for v_C=1:5
    for v_B=1:11
      numb_v=(v_B+11*(v_C-1))
      for bra=1:27
        for J1=1:150 # 150
          I_sharp[J1,bra]=Re[v_C,v_B]^2+*(1/(Meta[J1,bra,numb_v]-Meta_shift[numb_v])^4)*FCF[v_C,v_B]
          for k = K
            I_broad[k]+=I_sharp[J1,bra]*exp(-2*(Meta_cal[k]-(Meta[J1,bra,numb_v]-Meta_shift[numb_v]))^2/IB)
          end
        end 
      end
    end
  end
end
using LoopVectorization
function f2!(I_sharp, I_broad, Re, Meta, Meta_shift, FCF, Meta_cal, K, IB)
  @turbo for v_C=1:5
    for v_B=1:11
      numb_v=(v_B+11*(v_C-1))
      for bra=1:27
        for J1=1:150 # 150
          I_sharp[J1,bra]=Re[v_C,v_B]^2+*(1/(Meta[J1,bra,numb_v]-Meta_shift[numb_v])^4)*FCF[v_C,v_B]
          for k = K
            I_broad[k]+=I_sharp[J1,bra]*exp(-2*(Meta_cal[k]-(Meta[J1,bra,numb_v]-Meta_shift[numb_v]))^2/IB)
          end
        end 
      end
    end
  end
end

```

Note that `LoopVectorization` really wants `UnitRange` over `StepRange`, at least for `k`.

[Next page](https://discourse.julialang.org/t/question-improve-the-speed-of-for-loop/97204.md?page=2)
