# Map and mapreduce with Threads

**URL:** <https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861>\
**Category:** General Usage\
**Tags:** question\
**Created:** [December 4, 2019, 4:03pm UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861 "2019-12-04T16:03:15Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 4, 2019, 4:03pm UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/1 "2019-12-04T16:03:15Z")

</div>

I would like to experiment with parallelizing an embarrassingly parallel computation on an iterable using [the new multithreading in 1.3](https://julialang.org/blog/2019/07/multithreading). I need `map` and `mapreduce`.

I would prefer to leave return type determination to Julia, so I would prefer not to preallocate.

Do I just need to `@spawn` and then `fetch`, as in

```julia
using Base.Threads: @spawn, fetch, threadid, nthreads

@show nthreads()

function ploop(f, itr)
    map(fetch, map(i -> @spawn(f(i)), itr))
end

function pmapreduce(f, op, itr)
    mapreduce(fetch, op, map(i -> @spawn(f(i)), itr))
end

f(i) = (@show threadid(); sleep(rand()); Float64(i))

ploop(f, 1:10)
pmapreduce(f, +, 1:10)

```

(The motivation for this approach is that in practice `f` itself can use threads and the whole things hopefully composes neatly.)

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [December 4, 2019, 7:14pm UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/2 "2019-12-04T19:14:02Z")

</div>

This approach is taken in [https://github.com/baggepinnen/ThreadTools.jl](https://github.com/baggepinnen/ThreadTools.jl)

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [December 5, 2019, 7:12am UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/3 "2019-12-05T07:12:42Z")

</div>

For `mapreduce` it might make sense to split up the list recursively, so that `op` can also run in parallel.

---

<div class="post-metadata">

**Author:** ![longemen3000](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/longemen3000/32/7298_2.png) [@longemen3000](https://discourse.julialang.org/u/longemen3000)\
**Post date:** [December 5, 2019, 7:32am UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/4 "2019-12-05T07:32:59Z")

</div>

a robust `mapreduce` should include a lower threshold to activate. that threshold is related to the proportion of calling a function `f` and spawning a thread. for example, if` f(x) = x+1`, then the cost of spawning a thread is aproximately 5000 times more expensive (obviusly this is the worst case, but using a proportion between the time to do one operation and the time to spawn a thread can give a good approximation.

Spawning 1 thread per op seems excessive, but that strategy pays off if your operation is expensive enough to overcome the thread spawning overhead (by how much? 10 times?)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 5, 2019, 7:41am UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/5 "2019-12-05T07:41:45Z")

</div>

> [@longemen3000](#):
>
> a robust `mapreduce` should include a lower threshold to activate.

Obviously — the actual problem I am working on takes about 2–10 minutes for a single call, and the `sleep` in the MWE is a stand-in for this.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [December 5, 2019, 7:45am UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/6 "2019-12-05T07:45:22Z")

</div>

Then ThreadTools does exactly what you want.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [December 5, 2019, 9:39am UTC](https://discourse.julialang.org/t/map-and-mapreduce-with-threads/31861/7 "2019-12-05T09:39:22Z")

</div>

ThreadTools.jl is currently implemented with somewhat expensive computations in mind. There is a benchmark in the Readme that indicate roughly where the overhead starts becoming excessively expensive.
