# Julia: how to run embarrassingly parallel jobs with nested for loops?

**URL:** <https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410>\
**Category:** Julia at Scale\
**Tags:** parallel, multithreading\
**Created:** [July 10, 2021, 11:00am UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410 "2021-07-10T11:00:57Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![erwanlecarpentier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erwanlecarpentier/32/23448_2.png) [@erwanlecarpentier](https://discourse.julialang.org/u/erwanlecarpentier)\
**Post date:** [July 10, 2021, 11:00am UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/1 "2021-07-10T11:00:57Z")

</div>

Hey! I’d like to run 4 independent tasks in parallel for a certain number of times using 4 cores. Here is the script, say `test.jl`, that I run with the command `julia --threads 4 test.jl`:

```julia
using Distributed

function remote_computation()
    thid = Threads.threadid()
    println("Remote computation with thread $thid")
    sleep(5) # Heavy computation: to be parallelized
    # (usually I would put some real computation here but for reproducibility I used sleep)
    rand()
end

addprocs(Threads.nthreads()) # Add as many processes as available threads, 4 in my case

for t in 1:10 # Say we want to do the full parallelized process 10 times
    results_matrix = zeros(2,2)
    @sync for i in 1:2
        for j in 1:2
            @async Threads.@spawn results_matrix[i, j] = remote_computation()
        end
    end
    println("Completed parallelized process number: $t")
end

```

Notice that I have a nested for loop as, in my real application, `remote_computation` depends on `i` and `j` which explains why I have two levels of iteration. I removed this dependency for the sake of simplicity.

**Expectation:** My expectation is that for a sequence of 10 (outermost loop), the 4 processes are run in parallel and finish about at the same time. Also, if I do `htop`, I expect my 4 CPUs to be 100% busy during the whole execution (notice that in this MWE, using `sleep` may not make them busy, right?).  
In my understanding, that would give the following output.

```julia
Remote computation with thread 2
Remote computation with thread 3
Remote computation with thread 1
Remote computation with thread 4
Completed parallelized process number: 1
Remote computation with thread 1
Remote computation with thread 3
Remote computation with thread 2
Remote computation with thread 4
Completed parallelized process number: 2
etc... (10 times in total)

```

With the 4 CPUs 100% busy all along.

**Reality:** However, my expectation is only realized for the 1st of the 10 iterations and then jobs are almost run sequentially (only 2 threads active in parallel) and my cores are ~30% busy with one at 100%. One can see this with the fact that the same thread ID is being repeated starting from iteration 2.

```julia
Remote computation with thread 2
Remote computation with thread 3
Remote computation with thread 1
Remote computation with thread 4
Completed parallelized process number: 1
Remote computation with thread 1 # thread 1 used for 3 jobs out of 4
Remote computation with thread 1
Remote computation with thread 1
Remote computation with thread 4
Completed parallelized process number: 2
etc... (10 times in total)

```

I tried to follow [this discussion](https://stackoverflow.com/questions/37287020/how-and-when-to-use-async-and-sync-in-julia) to use `@sync` to have my script wait for the 4 processes are finished and `@async` to have them executed in parallel, but apparently I did miss something here.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [July 10, 2021, 1:17pm UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/2 "2021-07-10T13:17:15Z")

</div>

> [@erwanlecarpentier](#):
>
> `addprocs(Threads.nthreads())`

You’re mixing up threads and processes here. Are you planning to run multiple threads on multiple machines? If not, just stick to the threads and don’t spawn extra processes.

> [@erwanlecarpentier](#):
>
> ` @async`

Here you are mixing up threads and tasks. Read the docs on threads. Instead of calling @async and @sync you should be just spawning the threads and letting them run. If you want to synchronize them, use a channel or call wait or fetch.

---

<div class="post-metadata">

**Author:** ![erwanlecarpentier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erwanlecarpentier/32/23448_2.png) [@erwanlecarpentier](https://discourse.julialang.org/u/erwanlecarpentier)\
**Post date:** [July 10, 2021, 2:19pm UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/3 "2021-07-10T14:19:45Z")

</div>

Thanks for clarifying! I plan to use multiple threads on a **single** machine, thus I may stick to the threads as you recommended. I made a trial using `wait` with a “tasks matrix” but still observed the same “sequential” execution after iteration 2. Is it the right way to use wait?

```julia
for t in 1:10 # Say we want to do the full parallelized process 10 times
    results_matrix = zeros(2,2)
    tasks_matrix = Array{Task,2}(undef,2,2)
    for i in 1:2
        for j in 1:2
            tasks_matrix[i,j] = Threads.@spawn results_matrix[i, j] = remote_computation()
        end
    end
    for i in 1:2
        for j in 1:2
            wait(tasks_matrix[i,j])
        end
    end
    println("Completed parallelized process number: $t")
end

```

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [July 10, 2021, 4:13pm UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/4 "2021-07-10T16:13:53Z")

</div>

I think the fact that your remote\_computation calls sleep is problematic for your results. Change it as follows and see how your latest example works:

```julia
function remote_computation()
   thid = Threads.threadid()
   println("Remote computation with thread $thid")
   start = rand();
   for i in 1:10000000
      start += 1
   end
   return start
end

```

---

<div class="post-metadata">

**Author:** ![erwanlecarpentier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erwanlecarpentier/32/23448_2.png) [@erwanlecarpentier](https://discourse.julialang.org/u/erwanlecarpentier)\
**Post date:** [July 12, 2021, 6:07pm UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/5 "2021-07-12T18:07:19Z")

</div>

I think I got close to what I wanted using a combination of `Threads.@spawn` and `@sync` in the outer loop to wait for the results of each thread before continuing.

```julia
for t in 1:10 # Say we want to do the full parallelized process 10 times
    results_matrix = zeros(2,2)
    @sync for i in 1:2
        for j in 1:2
            Threads.@spawn begin
                results_matrix[i, j] = remote_computation()
            end
        end
    end
end

```

I do not observe the 4 CPUs at 100% though and still some threads are being used several times for the same parallelized iteration.

PS : I used a real remote computation function that does not use sleep as you suggested.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [July 12, 2021, 7:37pm UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/6 "2021-07-12T19:37:55Z")

</div>

> [@erwanlecarpentier](#):
>
> using a combination of `Threads.@spawn` and `@sync` in the outer loop

Yes, it looks like @sync handles threads

> " Wait until all lexically-enclosed uses of `@async` , `@spawn` , `@spawnat` and `@distributed` are complete. All exceptions thrown by enclosed async operations are collected and thrown as a `CompositeException` ."

So your solution should work right.

> [@erwanlecarpentier](#):
>
> I do not observe the 4 CPUs at 100% though

there is some overhead in spawning the thread and such, this can cause less than full utilization, particularly for tasks that take a small amount of time. If your task takes 5 seconds or more, you should see nearly full utilization. If your task takes 50ms or less you probably won’t. somewhere in between is the transition.

---

<div class="post-metadata">

**Author:** ![erwanlecarpentier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erwanlecarpentier/32/23448_2.png) [@erwanlecarpentier](https://discourse.julialang.org/u/erwanlecarpentier)\
**Post date:** [July 13, 2021, 9:22am UTC](https://discourse.julialang.org/t/julia-how-to-run-embarrassingly-parallel-jobs-with-nested-for-loops/64410/7 "2021-07-13T09:22:03Z")

</div>

Indeed I can see the overhead using a 20s task. I could see full utilization + equal distribution between the 4 threads on a remote machine dedicated to computing (no other processes running in the background), which is exactly what I was looking for.

On my own laptop I do not see that “perfect” distribution, but I guess this is because I am running other processes at the same time such as internet browser, IDE, etc.

Thanks a lot for your help!
