# Add rows to DataFrame (or similar table-like struct) in parallel?

**URL:** <https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721>\
**Category:** General Usage\
**Created:** [November 19, 2018, 5:31pm UTC](https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721 "2018-11-19T17:31:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![e3c6](https://avatars.discourse-cdn.com/v4/letter/e/e79b87/32.png) [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Post date:** [November 19, 2018, 5:31pm UTC](https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721/1 "2018-11-19T17:31:17Z")

</div>

I want to produce many rows of a `DataFrame` in parallel (using a `@distributed` for loop).

I like `DataFrames` because then I can `Query` it which is very convenient. But any other similar table-like structure that can be queried is good for me.

Any recommendations on how to do this? Specifically I need help on how to share the same table object across proceses that are writing to it.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 19, 2018, 5:51pm UTC](https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721/2 "2018-11-19T17:51:15Z")

</div>

I dont think that is currently feasible, but its on the radar for how this might work.

If you are adventurous, you could try forking DataFrames and changing the constructors to make all the arrays `SharedArray`s and let us know what problems you faced.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [November 19, 2018, 6:27pm UTC](https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721/3 "2018-11-19T18:27:08Z")

</div>

I use `vcat` as reducer for `@distributed`. This way all workers locally `vcat` their local parts of the overall `DataFrame`. Finally, the master (the process who `@distributed` the construction) `vcat`s the worker `DataFrame`s. The speed-up is sufficient for my use case. A `SharedDataFrame` would be great though.

---

<div class="post-metadata">

**Author:** ![e3c6](https://avatars.discourse-cdn.com/v4/letter/e/e79b87/32.png) [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Post date:** [November 19, 2018, 6:42pm UTC](https://discourse.julialang.org/t/add-rows-to-dataframe-or-similar-table-like-struct-in-parallel/17721/4 "2018-11-19T18:42:31Z")

</div>

I think this is enough for my purposes. Thanks!
