# Running simple loop in parallel

**URL:** <https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969>\
**Category:** New to Julia\
**Tags:** parallel\
**Created:** [June 26, 2018, 2:50pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969 "2018-06-26T14:50:06Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![myroslav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/myroslav/32/7354_2.png) [@myroslav](https://discourse.julialang.org/u/myroslav)\
**Post date:** [June 26, 2018, 2:50pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/1 "2018-06-26T14:50:06Z")

</div>

Hi,  
I am trying to do some simple computations inside a single or double loop, and I would like to parallelize the computations inside the loops, and so far I couldn’t figure it out. Here is a simple example of what I want to do

```julia

using Optim

const n_z = 5
const n_b = 5
const J = 60
const Ht_grid = [1.17; 1.50; 1.92]
const n_ht = length(Ht_grid)
const B_grid = [0.0; 0.00356596; 0.0285277; 0.096281; 0.228222]
const pension_grid = [0.280896; 0.309223; 0.340406; 0.374734; 0.412523]

const vvv_opt_r = zeros(n_z,n_b,J);
const policy_opt_r = zeros(2,n_z,n_b,J);

function test(x, y, z,w)
    return x+y+z+w
end

for i_z in 1:n_z, i_b in 1:n_b
    y = pension_grid[i_z]
    b = B_grid[i_b]
    bprime_opt_r = zeros(n_ht)
    vvv_r = zeros(n_ht)
    for i_ht = 1:n_ht
        htilde = Ht_grid[i_ht]
        fopt(x) = test(x,b,y,htilde)
        res = optimize(fopt, 0, maximum(B_grid))
        bprime_opt_r[i_ht] = Optim.minimizer(res)
        vvv_r[i_ht] = -Optim.minimum(res)
    end

    vvv_opt_r[i_z,i_b,J], idx = findmax(vvv_r)
    policy_opt_r[1,i_z,i_b,J] = bprime_opt_r[idx]
    policy_opt_r[2,i_z,i_b,J] = idx
end

```

Essentially the code loops over i\_z and i\_b, calculates something inside for each value of i\_z and i\_b (independtenly) and saves the resuls in vvv\_opt\_r, policy\_opt\_r and c\_opt\_r.  
All the functions (value\_age\_J\_r), arrays and constants (vvv\_opt\_r, policy\_opt\_r and c\_opt\_r, qb, current\_rent, etc.) are predifined in a different file that I load with “include(“path”)” command, as well as the packages (such as QuanEcon).

Ideally I would like to loop over both i\_z and i\_b in parallel and save the results to vvv\_opt\_r, etc as they come.

Thanks for the help,  
Myroslav

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [June 26, 2018, 2:51pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/2 "2018-06-26T14:51:02Z")

</div>

Please provide a minimal working example, and [quote your code](https://discourse.julialang.org/t/psa-how-to-quote-code-with-backticks/7530)

---

<div class="post-metadata">

**Author:** ![myroslav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/myroslav/32/7354_2.png) [@myroslav](https://discourse.julialang.org/u/myroslav)\
**Post date:** [June 26, 2018, 3:41pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/3 "2018-06-26T15:41:30Z")

</div>

done

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 26, 2018, 4:08pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/4 "2018-06-26T16:08:26Z")

</div>

Would simple @threads as described in the manual not work? It seems you wouldn’t run into race conditions/your code seems threadsafe.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [June 26, 2018, 4:31pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/5 "2018-06-26T16:31:10Z")

</div>

You should start by putting the main code in a function.  
You should be able to use DistributedArrays.jl for the parallelization.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 26, 2018, 7:12pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/6 "2018-06-26T19:12:37Z")

</div>

This is one of the domains that I am still very unclear on, maybe you can help me with this. At what point does it become better to use things like DistributedArrays and multiprocessing-style parallelism over shard-memory parallelism? Personally I haven’t found a usecase of these things yet, but admittedly I’m normally not dealing with gb’s - tb’s worth of data, nor have I found the need to parallelize over multiple cluster nodes yet.

But does it all just come down to data size, or the necessity to parallelize over nodes? A lot of the example code seems to be on semi trivial array sizes, and from personal experience shared memory threading seems to do a lot better in most cases. Has this changed drastically in the last half year? (last time I tried the multiprocessing stuff)

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [June 26, 2018, 9:55pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/7 "2018-06-26T21:55:16Z")

</div>

You can benchmark both. I’ve found it has changed as Julia has developed and also depends on the specific problem. I didn’t take (have) the time to figure out exactly why. But, it might be easy to test a simplified version of your problem.

---

<div class="post-metadata">

**Author:** ![myroslav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/myroslav/32/7354_2.png) [@myroslav](https://discourse.julialang.org/u/myroslav)\
**Post date:** [June 27, 2018, 10:18am UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/8 "2018-06-27T10:18:06Z")

</div>

Is there any chance you can show how it’s done? Either on my example or on a simpler one? Thanks

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 27, 2018, 2:39pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/9 "2018-06-27T14:39:59Z")

</div>

> [@myroslav](#):
>
> for i\_z in 1:n\_z, i\_b in 1:n\_b

```julia
Threads.@threads for i_z=1:n_z 
      for i_b=1:n_b 
           ...
       end
end

```

This should work, if I remember correct nested threading works. If not you might need to squash the two indices into 1.  
Oh and remember to run `export JULIA_NUM_THREADS=.... ` before launching julia

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [June 27, 2018, 2:52pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/10 "2018-06-27T14:52:15Z")

</div>

I’m not sure what that does.

But threading only the outer loop usually better. I tried threading on the OP’s code with `BenchmarkTools` and found no difference in speed using all four cores. You need something in the loop that uses more cpu time per iteration.

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [June 27, 2018, 2:56pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/11 "2018-06-27T14:56:06Z")

</div>

Wrap the code in a function so that you can benchmark it:

```julia
function runexample()
    Threads.@threads for i_z in 1:n_z
        for i_b in 1:n_b            
            y = pension_grid[i_z]
   ...

```

Then do

```julia
julia> using BenchmarkTools;
julia> @btime runexample()
  4.274 ms (202029 allocations: 3.12 MiB)

```

Try with and without threading, and threading on different loops.

EDIT: But make the arrays bigger. Also, you are allocating in the inner loop, which may be important. You can allocate the arrays outside the loops, one set for each thread and index them by threadid (look up how to get this in the docs). Then you can zero them before each iteration rather than reallocating. That takes time and adds complexity, so I’m not sure its worth it. You might want to try distributed parallelism first.

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 27, 2018, 2:56pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/12 "2018-06-27T14:56:49Z")

</div>

It doesn’t work or it doesn’t provide any speedup. I assume that in the actual code a bit more is happening than in the MWE? I agree on the statement that you should usually only thread outer loop, I just saw that it were only 5 iterations so I thought that maybe doing the full 10 would be at least a bit more reason to thread 😃.

On the outer loop threading, in the situation that one runs a lot of matrix operations that rely on BLAS, it might be better to lower the amount of outer threads in favor of the amount of threads used for BLAS ( I think `BLAS.set_num_threads(nthreads)` is the way to set those). Some testing on the best ratio might result in minor speedups. Also, when using the full amount of threads for the outer loop, be sure to put the BLAS threads to 1, otherwise they might interfere and severely impact the performance (happened to me a couple of times).

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [June 27, 2018, 3:03pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/13 "2018-06-27T15:03:33Z")

</div>

Threading on the outer loop runs the inner loop in separate threads “automatically” so-to-speak, just because the entire inner loop is running in separate threads.

EDIT:  
If you have enough memory for multiple copies of data that varies with iterations, it is better to disable BLAS threading and thread at the outermost loop possible. This is general advice from my experience. But, it depends on details of your problem. If you are operating on huge matrices with builtin functions that consume significant time, then threading only via BLAS might be better.

Ha discourse is showing your response before I am done. I see your point. But, it would take more work to get 25 threads (and btw, this involves yet another parameter, how many cores do you have have ? and will hyperthreading work well, or not. This may depend on the chipset and details of the computation)

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 27, 2018, 3:09pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/14 "2018-06-27T15:09:38Z")

</div>

Right that’s true. The way I wrote it would mean that the inner loop itself would also get multithreaded. Imagine that the code runs on a 25 core machine, instead of only using 5 threads (since in the MWE `n_z = 5`) this would mean that each of the 5 threads spawns another 5 to solve the inner loop.

But I just checked the `base/threadingconstructs.jl` and it turns out that’s not what would happen so let me change my post. 😄

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [June 27, 2018, 3:25pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/15 "2018-06-27T15:25:13Z")

</div>

Also, you are violating the tip [Access arrays in memory order, along columns](https://docs.julialang.org/en/v0.6.1/manual/performance-tips/#Access-arrays-in-memory-order,-along-columns-1). Although in this case, it probably will not make any difference.

---

<div class="post-metadata">

**Author:** ![myroslav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/myroslav/32/7354_2.png) [@myroslav](https://discourse.julialang.org/u/myroslav)\
**Post date:** [June 27, 2018, 4:44pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/16 "2018-06-27T16:44:53Z")

</div>

thanks for the advice guys, I will try with the benchmarking and see what it does.  
and yes, my actual code will do more computationally intensive stuff inside the loop, so it might help speed up the overall code quite a bit!

---

<div class="post-metadata">

**Author:** ![myroslav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/myroslav/32/7354_2.png) [@myroslav](https://discourse.julialang.org/u/myroslav)\
**Post date:** [June 27, 2018, 4:52pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/17 "2018-06-27T16:52:05Z")

</div>

and regarding the number of threads, it’s at most equal to number of cores, right? there is no way I can run 8 threads at the same time on a 4 core machine?

---

<div class="post-metadata">

**Author:** ![saschatimme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saschatimme/32/10313_2.png) [@saschatimme](https://discourse.julialang.org/u/saschatimme)\
**Post date:** [June 27, 2018, 6:09pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/18 "2018-06-27T18:09:22Z")

</div>

> [@myroslav](#):
>
> there is no way I can run 8 threads at the same time on a 4 core machine?

You can, but this would completely trash the performance of your code (unless your CPU has hyperthreading but even this would probably not yield any improvements).

---

<div class="post-metadata">

**Author:** ![louisponet](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louisponet/32/2070_2.png) [@louisponet](https://discourse.julialang.org/u/louisponet)\
**Post date:** [June 27, 2018, 7:51pm UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/19 "2018-06-27T19:51:36Z")

</div>

I find that hyperthreading does lead to some improvements for most things. If it’s all BLAS/hardcore optimized math then no, but usually that doesn’t happen. I tend to see between 10-30% improvement.

---

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [July 3, 2018, 8:44am UTC](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969/20 "2018-07-03T08:44:24Z")

</div>

I had a similar question, and wrote up the answers (various parallelizing choices) in

[http://cookbook.tips/!julia/doku.php?id=parallel](http://cookbook.tips/!julia/doku.php?id=parallel)

hope this helps.

regards,

/iaw

[Next page](https://discourse.julialang.org/t/running-simple-loop-in-parallel/11969.md?page=2)
