# Parallel Postprocessing

**URL:** <https://discourse.julialang.org/t/parallel-postprocessing/29487>\
**Category:** Julia at Scale\
**Tags:** parallel\
**Created:** [October 4, 2019, 3:19pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487 "2019-10-04T15:19:29Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [October 4, 2019, 3:19pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/1 "2019-10-04T15:19:29Z")

</div>

Hi everyone,

I am using Julia PyPlot to do postprocessing for simulation data. Here is the sample script for my parallel work:

```julia
using Distributed
@everywhere using PyCall
@everywhere matplotlib = pyimport("matplotlib")
@everywhere matplotlib.use("Agg")
@everywhere using PyPlot, Glob

@everywhere function process(filename::String, dir::String=".")
   np = pyimport("numpy");
   filehead, data, filelist = readdata(filename, dir=dir, verbose=false);
   # Postprocessing...
   plt.savefig("$(time).png")
   println("finished saving $(time)!")
end

# Define path and filenames
dir = ".";
filename = "y*.out";

# Find filenames
# ......

# Processing
@distributed for filename in filenames
    println("filename: $(filename)")
    process(filename, dir)
end

```

In this way, I find all the filenames on one processor, distribute the names within workers, and do the plotting and saving on each worker using @distributed. Since there’s no dependency between processing different files, this seems to work.

I am wondering if there are better ways to do this, say, using @threads or channel? Any idea is appreciated!

---

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [October 4, 2019, 4:33pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/2 "2019-10-04T16:33:19Z")

</div>

For embarrassingly parallel tasks like this, this looks good to me. Simple and effective.

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [October 4, 2019, 8:06pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/3 "2019-10-04T20:06:52Z")

</div>

Later I encountered some issue when running this script in the command line but not in REPL. In REPL mode, everything looks fine, the plots are saved in png format. However, if I just type

```julia
julia process.jl

```

Then the function process is never been executed, and no error message returned. What’s wrong with that? I feel like the scheduled task in the queue is never been executed.

---

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [October 4, 2019, 9:42pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/4 "2019-10-04T21:42:18Z")

</div>

The driver script is sending the tasks off to workers and exiting immediately, because the distribute for loop returns immediately. This in turn kills the workers immediately. Doesn’t happen on the REPL because that stays alive after you run each command.

You need to have your driver script wait for the workers to all finish, by for example doing a dummy reduce.

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [October 4, 2019, 9:50pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/5 "2019-10-04T21:50:17Z")

</div>

Makes sense! So what is the command I’m missing? Some synchronization, barrier() or just wait()?

---

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [October 4, 2019, 10:06pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/6 "2019-10-04T22:06:56Z")

</div>

The docs for @distributed say that if you give it a reducer it’ll wait for the workers, so that it can compute the reduction. So you can have each worker return something, say 1, and use + as the reducer.

Or just add @sync in front of @distributed.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [October 4, 2019, 11:54pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/7 "2019-10-04T23:54:17Z")

</div>

I created ThreadTools for this kind of tasks  
[https://github.com/baggepinnen/ThreadTools.jl](https://github.com/baggepinnen/ThreadTools.jl)

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [October 5, 2019, 2:27am UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/8 "2019-10-05T02:27:25Z")

</div>

Is this for the upcoming 1.3 only? I would think thread pools is preferable in these kind of parallel IO work.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [October 5, 2019, 5:52am UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/9 "2019-10-05T05:52:31Z")

</div>

Yeah it only works for v1.3  
I have a small benchmark in the Readme indicating roughly where the overhead caused by my implementation strategy starts becoming noticeable.

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [October 5, 2019, 2:17pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/10 "2019-10-05T14:17:45Z")

</div>

Can’t wait to try v1.3 and your package! Thanks for sharing!

---

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [October 5, 2019, 7:00pm UTC](https://discourse.julialang.org/t/parallel-postprocessing/29487/11 "2019-10-05T19:00:21Z")

</div>

Nice!

@baggepinnen your response promoted me to catch up on the [multithreading news](https://julialang.org/blog/2019/07/multithreading)

If I understand the single global lock on libuv correctly, that means that interacting with it is done on a privileged thread, but the actual IO can happen in parallel? If not then for IO heavy tasks threads wouldn’t help…
