# What's the most efficient way to sum the results of multithreading?

**URL:** <https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944>\
**Category:** Performance\
**Tags:** multithreading, threads\
**Created:** [May 27, 2021, 3:41pm UTC](https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944 "2021-05-27T15:41:18Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![yingqiuz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yingqiuz/32/25165_2.png) [@yingqiuz](https://discourse.julialang.org/u/yingqiuz)\
**Post date:** [May 27, 2021, 3:41pm UTC](https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944/1 "2021-05-27T15:41:18Z")

</div>

Hi,

I need to sum the outputs of each thread. Is there a standard/efficient way to do this?  
for example:

```julia
X = Matrix{Float64}(d, N)
Threads.@threads for k = 1:N
    X[:, k] .= some_calculations()
end
sum!(ones(Int64, d), X)

```

Preferably avoid initialising an X, because I will only need the row sum.  
Sorry about it if this is a stupid question. Any help would be much appreciated!

---

<div class="post-metadata">

**Author:** ![Satvik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/satvik/32/20486_2.png) [@Satvik](https://discourse.julialang.org/u/Satvik)\
**Post date:** [May 27, 2021, 3:56pm UTC](https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944/2 "2021-05-27T15:56:26Z")

</div>

I think the easiest way is to use `ThreadsX.mapreduce`: [https://github.com/tkf/ThreadsX.jl](https://github.com/tkf/ThreadsX.jl)

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 28, 2021, 2:26am UTC](https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944/3 "2021-05-28T02:26:41Z")

</div>

ThreadsX.jl or any JuliaFolds-related packages do not have out-of-the-box multi-dimensional reduction API. This is mainly because Julia already has a rich set of DSL packages such as [Tullio.jl](https://github.com/mcabbott/Tullio.jl), [TensorOperations.jl](https://github.com/Jutho/TensorOperations.jl), and [LoopVectorization.jl](https://github.com/JuliaSIMD/LoopVectorization.jl) etc. to support it.

Having said that, you can use [`Broadcasting` transducer](https://juliafolds.github.io/Transducers.jl/dev/reference/manual/#Transducers.Broadcasting) to construct a custom multi-dimensional reduction:

```julia
julia> using Folds, Transducers, Statistics

julia> randn(2, 3) |> eachcol |> Map(x -> x ./ mean(x)) |> Broadcasting() |> Folds.sum
2-element Vector{Float64}:
 -3.962371501856117
  9.962371501856117

```

or simply looping over the input in the other way around:

```julia
julia> Folds.collect(sum(xs) for xs in eachcol(randn(2, 3)))
3-element Vector{Float64}:
 -1.9645612346392887
  0.16303379291104828
 -0.22374280710495784

```

---

<div class="post-metadata">

**Author:** ![yingqiuz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yingqiuz/32/25165_2.png) [@yingqiuz](https://discourse.julialang.org/u/yingqiuz)\
**Post date:** [May 28, 2021, 10:23pm UTC](https://discourse.julialang.org/t/whats-the-most-efficient-way-to-sum-the-results-of-multithreading/61944/4 "2021-05-28T22:23:53Z")

</div>

> [@tkf](#):
>
> `|> Broadcasting() |> Folds.sum`

Hi, thank you for your reply. LoopVectorization works for me.  
I also tried Broadcasting transducer, but it was not as good as ThreadsX, even with allocations. Could you shed some light on this?

```julia
julia> @btime ThreadsX.sum(x -> (sleep(0.5); x .* transpose(x)), eachcol(randn(1000, 50)))  
  2.106 s (524 allocations: 755.72 MiB)

julia> @btime randn(1000, 50) |> eachcol |> Map(x -> (sleep(0.5); x .* transpose(x))) |> Broadcasting() |> Folds.sum
  12.717 s (358 allocations: 397.12 MiB)

```

Thank you very much!
