# What is the “standard” way to use multi-threading with dot-calls (vectorized/broadcast)?

**URL:** <https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473>\
**Category:** General Usage\
**Tags:** multithreading\
**Created:** [April 24, 2024, 7:10pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473 "2024-04-24T19:10:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![lmtzx9h4qqnt](https://avatars.discourse-cdn.com/v4/letter/l/74df32/32.png) [@lmtzx9h4qqnt](https://discourse.julialang.org/u/lmtzx9h4qqnt)\
**Post date:** [April 24, 2024, 7:10pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/1 "2024-04-24T19:10:48Z")

</div>

All the discussions I’ve found on this topic are years old:

- [“Multithreaded broadcast?” thread](https://discourse.julialang.org/t/multithreaded-broadcast/26786)
- [use multiple threads in vectorized operations · Issue #1802 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/1802)
- [multi-threaded (@threads) dotcall/broadcast? · Issue #19777 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/19777)
- [`@threads f(...)` syntax to tell callee to use threaded version? · Issue #40329 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/40329)

What is currently the most straight-forward way to get multi-threading with operations like these:

```julia
a = rand(4000, 4000)
b = rand(4000, 4000)
c = zeros(4000, 4000)
c .= a .+ exp.(b)

```

Does everyone just use [Strided.jl](https://github.com/Jutho/Strided.jl)? (I think [LoopVectorization.jl](https://github.com/JuliaSIMD/LoopVectorization.jl) also had some kind of threading capability, but the package is now deprecated, and either way, not as general and much more invasive than simply parallelizing independent operations across threads OpenMP-style.)

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [April 24, 2024, 7:14pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/2 "2024-04-24T19:14:37Z")

</div>

I use the multithreaded map functions from OhMyThreads.jl or ThreadsX.jl.

---

<div class="post-metadata">

**Author:** ![lmtzx9h4qqnt](https://avatars.discourse-cdn.com/v4/letter/l/74df32/32.png) [@lmtzx9h4qqnt](https://discourse.julialang.org/u/lmtzx9h4qqnt)\
**Post date:** [April 24, 2024, 7:15pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/3 "2024-04-24T19:15:59Z")

</div>

Neither of those work with broadcasts though, as far as I can tell?

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [April 24, 2024, 8:19pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/4 "2024-04-24T20:19:31Z")

</div>

Right, `map` is an alternative to broadcast.

---

<div class="post-metadata">

**Author:** ![PeterSimon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petersimon/32/25193_2.png) [@PeterSimon](https://discourse.julialang.org/u/PeterSimon)\
**Post date:** [April 25, 2024, 11:30pm UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/5 "2024-04-25T23:30:38Z")

</div>

> [@lmtzx9h4qqnt](#):
>
> Does everyone just use [Strided.jl](https://github.com/Jutho/Strided.jl)

I hadn’t heard of this amazing package until you mentioned it. For fun, and to practice my almost-forgotten macro-fu, I made an attempt to write a short macro to save some typing when combining `@strided` and `@.`:

```julia
module StridedDot

export @sd

using Strided: Strided, _strided, maybestrided, sreshape, sview, maybeunstrided
using MacroTools: @capture, postwalk

function xform(ex)
    postwalk(ex) do x
        @capture(x, Strided) || return x
        return :StridedDot
    end
end

macro sd(ex1)
    ex = Strided._strided(Base.Broadcast. __dot__ (ex1))
    esc(xform(ex))
end

end # module

```

Then the timing comparison on my 8-core Core i7-9700 is:

```julia
using BenchmarkTools

a = rand(4000, 4000)
b = rand(4000, 4000)
c = zeros(4000, 4000)

f1!(c, a, b) = @. c = a + exp(b)

@btime f1!($c, $a, $b) # 63.204 ms (0 allocations: 0 bytes)

using .StridedDot
f2!(c, a, b) = @sd c = a + exp(b)
c2 = similar(c)
@btime f2!($c2, $a, $b) # 15.843 ms (125 allocations: 11.42 KiB)

c == c2 # true

```

---

<div class="post-metadata">

**Author:** ![lmtzx9h4qqnt](https://avatars.discourse-cdn.com/v4/letter/l/74df32/32.png) [@lmtzx9h4qqnt](https://discourse.julialang.org/u/lmtzx9h4qqnt)\
**Post date:** [April 26, 2024, 4:47am UTC](https://discourse.julialang.org/t/what-is-the-standard-way-to-use-multi-threading-with-dot-calls-vectorized-broadcast/113473/6 "2024-04-26T04:47:10Z")

</div>

Unfortunately, I just noticed the following discussion in [Correct way to parallelize this code? · Issue #9 · Jutho/Strided.jl · GitHub](https://github.com/Jutho/Strided.jl/issues/9)

> [@Jutho (Strided.jl author)](#):
>
> The simple case of a parallelising a plain broadcast operation was not my main concern when implementing this package. Maybe that particular case can be separated out early in the analysis, but probably there are already other packages which are much better at this.

> [@johnomotani](#):
>
> 👍 thanks for the info!
> 
> > probably there are already other packages which are much better at this.
> 
> Sorry, I’m very new to Julia… if it’s obvious what any of those other packages are, I’d be interested to know.

> [@Jutho (Strided.jl author)](#):
>
> I have not been following very actively myself all the recent developments. I think you want to check out things like LoopVectorization.jl , which will also do other optimizations

And again the discussion ends around the same time frame, 2021. Not that it’s not a good option in those cases where it works, but maybe not exactly the ideal candidate for a “standard” solution.
