# Processing Rows with multiple threads

**URL:** <https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253>\
**Category:** Performance\
**Tags:** question\
**Created:** [August 20, 2020, 8:32am UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253 "2020-08-20T08:32:19Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [August 20, 2020, 8:32am UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/1 "2020-08-20T08:32:19Z")

</div>

I find myself having to process a lot of data recently where a) the processing of two different rows is independent and b) processing each row is expensive. I find myself coding a version of the following every time:

```julia
using DataFrames
using Base.Threads: @spawn
myrows = rand(20)
function processmyrow(row) # => some expensive function
    sleep(5)
    return DataFrame(r = row)
end
outdfs = [DataFrame() for _ in 1:Threads.nthreads()]
myparts = Iterators.partition(myrows, cld(length(myrows), Threads.nthreads()))

@sync for part in myparts
    @spawn begin
        for row in part
            append!(outdfs[Threads.threadid()], processmyrow(row))
        end
    end
end

```

Is there a better way of doing this? I have started looking into a couple of parallelism exploiting packages recently its quite hard for me to comprehend. Could someone point me in the right direction here?

Edit: I do not care about the order of the result as I will usually just sort it afterwards.

Thanks!

---

<div class="post-metadata">

**Author:** ![haberdashPI](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/haberdashpi/32/26337_2.png) [@haberdashPI](https://discourse.julialang.org/u/haberdashPI)\
**Post date:** [August 20, 2020, 12:11pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/2 "2020-08-20T12:11:38Z")

</div>

> [@danielw2904](#):
>
> ```julia
> outdfs = [DataFrame() for _ in 1:Threads.nthreads()]
> myparts = Iterators.partition(myrows, cld(length(myrows), Threads.nthreads()))
> 
> @sync for part in myparts
> @spawn begin
> for row in part
> append!(outdfs[Threads.threadid()], processmyrow(row))
> end
> end
> end
> 
> ```

Instead of the above, you can do the following.

```julia
using BangBang, Transducers
foldxt(append!!, Map(processmyrow), myrow))

```

The doc for Transducers has a good [parallelism tutorial](https://juliafolds.github.io/Transducers.jl/stable/tutorials/tutorial_parallel)

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [August 20, 2020, 12:23pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/3 "2020-08-20T12:23:09Z")

</div>

Thank you that looks great!

I also found

```julia
tcollect(Map(processmyrow), myrows))

```

in case the rows cannot easily be appended (e.g. not all rows return the same collumns)

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [August 20, 2020, 12:25pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/4 "2020-08-20T12:25:24Z")

</div>

> [@haberdashPI](#):
>
> The doc for Transducers has a good [parallelism tutorial](https://juliafolds.github.io/Transducers.jl/stable/tutorials/tutorial_parallel/https://juliafolds.github.io/Transducers.jl/stable/tutorials/tutorial_parallel/).

I think you pasted the link twice its: [parallelism tutorial](https://juliafolds.github.io/Transducers.jl/stable/tutorials/tutorial_parallel/)

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [August 20, 2020, 12:36pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/5 "2020-08-20T12:36:39Z")

</div>

If the function on a row takes sufficently long time (e.g. 1 ms or more), you should be fine to spawn a task per row and construct a DataFrame from the result:

```julia
df = DataFrame(a=1:10, b=11:20)
processmyrow(row) = (a=row.a, b=row.b, c=row.a^2 + row.b)
outdf = DataFrame(fetch.([Threads.@spawn processmyrow(row) for row in eachrow(df)]))

```

---

<div class="post-metadata">

**Author:** ![haberdashPI](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/haberdashpi/32/26337_2.png) [@haberdashPI](https://discourse.julialang.org/u/haberdashPI)\
**Post date:** [August 20, 2020, 12:49pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/6 "2020-08-20T12:49:59Z")

</div>

> [@danielw2904](#):
>
> I think you pasted the link twice its:

Bah! That’s what I get for typing with a baby in one arm…

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [August 20, 2020, 2:36pm UTC](https://discourse.julialang.org/t/processing-rows-with-multiple-threads/45253/7 "2020-08-20T14:36:09Z")

</div>

Transducers has changed my life! I already added it to 2 projects I’m working on.
