# Recommended way to do work stealing in Julia 1.0?

**URL:** <https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573>\
**Category:** Julia at Scale\
**Tags:** question, performance\
**Created:** [September 27, 2018, 4:51pm UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573 "2018-09-27T16:51:39Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![orenbenkiki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/orenbenkiki/32/5578_2.png) [@orenbenkiki](https://discourse.julialang.org/u/orenbenkiki)\
**Post date:** [September 27, 2018, 4:51pm UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/1 "2018-09-27T16:51:39Z")

</div>

I have a Python system which is a mixture of C++ code, pandas/numpy code running on top of BLAS/MKL, and plain Python code binding the whole thing together. “It ain’t a pretty sight”.

It can be viewed as a nested tree (actually DAG) of tasks, with some tasks invoking several dozen sub-tasks, with a total of hundreds of tasks, which may grow to thousands as the project develops.

I got this to work with Python multi-processing by picking a specific level in the tasks tree to distribute across multiple CPUs, but this is sub-optimal as there is variance between the execution times of the sub-trees. Also getting shared-memory Pandas data frames is possible, but very fragile.

It is obvious I’m pushing the platform beyond what it is intended for… So I am looking for alternatives. I have been tracking Julia for a while and now it as just released v1.0 it sounded like it might be a practical choice. The libraries I make use of mostly do basic linear algebra stuff, nothing fancy, so porting shouldn’t be that hard. Using a single language instead of having to dance between my python code, pandas/numpy, and custom C++ extensions sounds really tempting.

What I need is almost exactly is described in [Shared memory parallelism in Julia with multi-threading | Cambridge Julia Meetup (May 2018) - YouTube](https://www.youtube.com/watch?v=YdiZa0Y3F3c) - however, looking at [WIP: parallel task runtime by kpamnany · Pull Request #22631 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/pull/22631) I see this didn’t make it to v1.0. However, reading through [Parallel Computing · The Julia Language](https://docs.julialang.org/en/v1/manual/parallel-computing/index.html) I see that the low-level building blocks seem to exist (I’m not interested in async I/O issues - my project is almost pure computations).

So: What is the current Julia v1.0 best practice for writing a system that tries to compute a DAG of compute tasks across multiple threads, using shared memory between the tasks for optimal performance, with depth-first work stealing?

In particular, should I:

- Stick with my Python system (which mostly works) until PR22631 is merged into a Julia point release. This would be reasonable for my project if it happens in the next, say, 1 or 2 quarters. If it wouldn’t be merged within, say, a year, then just waiting for it wouldn’t make sense.

- Roll my own system? I understand this was done in [https://juliaimages.github.io/latest/imagefiltering.html](https://juliaimages.github.io/latest/imagefiltering.html) for example. I only expect hundreds or thousands of not-too-fast tasks so a naive thread-safe priority queue might not be a bottleneck for me. And there are probably other relaxations I can take advantage of with a tailored system. But if Julia will provide a built-in solution “soon”, this work would be a pure waste.

- Use some package that provides a similar functionality? Come to think of it, it isn’t clear to me why PR22631 is a PR instead of a package building on top of the mechanisms already provided by Julia v1.0. This is not to say I don’t appreciate such a generic, high-performance production-quality package would be far from trivial. It is certainly beyond the scope of what I can afford creating within my own project.

- Some other option?

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [September 27, 2018, 8:46pm UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/2 "2018-09-27T20:46:19Z")

</div>

Maybe not exactly what you are looking for (as it uses distributed memory) but there is a Julia package to run computations represented as directed acyclic graphs efficiently: [GitHub - JuliaParallel/Dagger.jl: A framework for out-of-core and parallel execution](https://github.com/JuliaParallel/Dagger.jl)

---

<div class="post-metadata">

**Author:** ![orenbenkiki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/orenbenkiki/32/5578_2.png) [@orenbenkiki](https://discourse.julialang.org/u/orenbenkiki)\
**Post date:** [September 28, 2018, 7:05am UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/3 "2018-09-28T07:05:25Z")

</div>

Thanks, Dagger looks nice! I might use it for distributing coarse-grained work across a cluster. The lack of just-shared-memory restricts its applicability for me, though. I have large-ish data sets (100s of MBs) which don’t have a natural partition into shards, so are much better suited for a shared rather than a distributed memory implementation. And I have plenty of machines with \>50 threads to play with…

---

<div class="post-metadata">

**Author:** ![Zach\_Christensen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zach_christensen/32/7220_2.png) [@Zach\_Christensen](https://discourse.julialang.org/u/Zach_Christensen)\
**Post date:** [September 28, 2018, 11:51am UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/4 "2018-09-28T11:51:49Z")

</div>

It would be nice if there were something like Apache-airflow or Nextflow in Julia. It looks like JuliaRun is probably comparable, but it’s hard to justify paying for something like that when there are these other free alternatives.

---

<div class="post-metadata">

**Author:** ![orenbenkiki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/orenbenkiki/32/5578_2.png) [@orenbenkiki](https://discourse.julialang.org/u/orenbenkiki)\
**Post date:** [September 29, 2018, 1:51pm UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/5 "2018-09-29T13:51:57Z")

</div>

JuliaRun seems to be an overkill for what I’m looking for… But it is an interesting option if I want to scale things to a cluster setting - and if they have some sort of free licenses for academic projects.

---

<div class="post-metadata">

**Author:** ![kpamnany](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kpamnany/32/206168_2.png) [@kpamnany](https://discourse.julialang.org/u/kpamnany)\
**Post date:** [November 20, 2018, 3:29am UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/6 "2018-11-20T03:29:24Z")

</div>

PR22631 (aka `partr`) will almost certainly be merged before the end of the year. However it will not be the default task runtime for some time after that. There is still a good bit of work left to do before parallel tasks are a well tested and stable thing, and broadly available in a binary release. Your use case would be helpful to that end, if you can share the code together with validation logic and are willing to help us debug. 🙂

Edit: whoops! I didn’t notice the date on this one.

---

<div class="post-metadata">

**Author:** ![jobjob](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jobjob/32/160_2.png) [@jobjob](https://discourse.julialang.org/u/jobjob)\
**Post date:** [February 20, 2019, 1:48pm UTC](https://discourse.julialang.org/t/recommended-way-to-do-work-stealing-in-julia-1-0/15573/7 "2019-02-20T13:48:02Z")

</div>

@kpamnany any updates on when 22631 is likely to get merged? Basically the only thing left is weird bugs on Windows?
