# Multi-threading for nested loops with inner loop collection depending on outer loop

**URL:** <https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569>\
**Category:** Performance\
**Tags:** parallel, multithreading, io\
**Created:** [April 17, 2023, 12:44pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569 "2023-04-17T12:44:12Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [April 17, 2023, 12:44pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/1 "2023-04-17T12:44:12Z")

</div>

What is the most efficient way to parallelize an embarrassingly parallel problem with nested loops, while the inner loop’s collection depending on outer loop? MWE is

```julia
using Base.Threads
using ThreadsX
using FLoops

xs = 4:8
ys = Dict(xs .=> [1:rand(1:3) for _ in xs])

for x in xs
    for y in ys[x]
        # f(x, y) # an expensive function most of its time for file IO and database IO jobs.   
        @show threadid(), x, y
    end
end

```

The line `f(x, y)` is an expensive function. The workload of `f` is highly unbalanced, which depends on specific (x, y). What `f` does is downloading data form ClickHouse, processing them, and then save to local CSV files.

It seems I cannot use `ThreadsX.foreach` or `ThreadsX.map` directly.

Also there are many ways to put `@threads` or `@spawn` or `@floop` before two `for`s. Which one is best?

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [April 17, 2023, 2:51pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/2 "2023-04-17T14:51:45Z")

</div>

I’ve been playing with

```julia
using Base.Threads

xs = 4:8
ys = Dict(xs .=> [1:rand(1:3) for _ in xs])

@threads for (x, y) in [(x, y) for x in xs for y in ys[x]]
    # f(x, y) # an expensive function most of its time for file IO and database IO jobs.   
    @show threadid(), x, y
end

```

Not sure if this helps with your real task though.

---

<div class="post-metadata">

**Author:** ![cgeoga](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cgeoga/32/216186_2.png) [@cgeoga](https://discourse.julialang.org/u/cgeoga)\
**Post date:** [April 17, 2023, 6:37pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/3 "2023-04-17T18:37:40Z")

</div>

I don’t know if this works for your problem, but one lazy option to consider is just pre-computing `(x,y)` and then doing a simple loop:

```julia
xypairs = reduce(vcat, map(x->[(x, yj) for yj in ys[x]], xs)) # or whatever
@threads for (x,y) in xypairs
  #[...]
end

```

maybe that isn’t possible, but whatever overhead that incurs might be more than made up with better parallel execution in the loop.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [April 17, 2023, 6:45pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/4 "2023-04-17T18:45:02Z")

</div>

> [@liuyxpp](#):
>
> `xs = 4:8`

Is this range realistic (only 4 steps)? If so, merging the loops is probably the way to go, one way or the other. If `xs` length is large then the story might be different.

---

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [April 18, 2023, 1:30am UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/5 "2023-04-18T01:30:41Z")

</div>

This is a MWE. The actual length of xs should be around 200-1000, and for each ys its length is about 2000-5000. `f(x,y)` roughly takes 500-1000 ms, when x, y points missing data, it immediately returns (about 1 ms).

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [April 18, 2023, 12:11pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/6 "2023-04-18T12:11:33Z")

</div>

If there is no concurrency, I think the most effective way is to use `@spawn`, like this:

```julia
Threads.@sync for x in xs
    for y in ys[x]
        Threads.@spawn f(x, y) # an expensive function most of its time for file IO and database IO jobs.   
    end
end

```

which will run `f(x,y)` on any free thread available.

---

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [April 18, 2023, 12:22pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/7 "2023-04-18T12:22:53Z")

</div>

Yes, my experiment confirms it. However, the CPU loading is oscillating between 10% to 80% every 1-10 seconds. I have already set number of threads to 32 which is 2x number of cores of my machine. Maybe it is caused by crazy number of allocations? I notice the GC time occupies upto 50%. And the size of memory allocated is upto 1T.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [April 18, 2023, 12:25pm UTC](https://discourse.julialang.org/t/multi-threading-for-nested-loops-with-inner-loop-collection-depending-on-outer-loop/97569/8 "2023-04-18T12:25:25Z")

</div>

It may be, but also if what the tasks do is to write stuff to the disc, you may have a bottleneck in the concurrency for the disk resources. (but in general, avoiding allocations also allows threading code to run faster, yes)
