# Propagation of available/assigned worker-IDs in hierarchical computations?

**URL:** <https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670>\
**Category:** Julia at Scale\
**Created:** [May 3, 2020, 1:06pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670 "2020-05-03T13:06:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 3, 2020, 1:06pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/1 "2020-05-03T13:06:21Z")

</div>

I’ve been thinking about ways to propagate information about available workers (and possibly other resources) in scenarios with nested computations. Say we have 10000 workers available and want to run a high-level distributed algorithm which will scale up to 100 workers. Each parallel part of the high-level algorithm will in turn use a lower-level parallel algorithm internally that can also scale up to 100 workers. So all 10000 workers could be utilized - the question is, how will each instance of the low-level algorithm know _which_ workers it’s allowed to use, when coding this in a modular fashion: There could be different low-level algorithms to choose from, and there might also be methods that are used stand-alone in other situations - we don’t want to code it as a monolithic thing.

On the thread-level, the partr scheduler has pretty much solved these problem now (since Julia v1.3). To my knowledge, we currently don’t have such a scheduler for (possibly distributed) worker processes. I was thinking about some kind of simple solution we could use until we have a fancy scheduler for workers, like we have for threads now.

Maybe `task_local_storage` could be used to propagate information about available/assigned resources (mainly workers) in a hierarchy of tasks? We’d need a way to pass it along when spawning new local/remote tasks, of course.

If so, could we come up with a community standard on which key names/values to use in `task_local_storage`, so that we can pass resource availability information through different parallel computing packages (e.g. [Transducers.jl](https://github.com/tkf/Transducers.jl), CC @tkf) in a hierarchical computation?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 3, 2020, 3:14pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/2 "2020-05-03T15:14:09Z")

</div>

It’d be _awesome_ if you can come up with partr-like abstraction for distributed computation! Naively thinking, some ideas behind partr like worker-local queues (work stealing) can work in distributed setting? I’m too noob to guess anything about it ATM, though.

> [@oschulz](#):
>
> Maybe `task_local_storage` could be used to propagate information about available/assigned resources (mainly workers) in a hierarchy of tasks?

IIUC, this doesn’t work because `task_local_storage` is not propagated to sub-tasks:

```julia
julia> @sync @async begin
           task_local_storage(:mykey, "hello")
           @show task_local_storage(:mykey)
           @async try
               @show task_local_storage(:mykey)
           catch err
               @show err
           end
       end;
task_local_storage(:mykey) = "hello"
err = KeyError(:mykey)

```

I think we need something like `task_local_storage` but that are inherited by sub-tasks do this (ref [Context Variables in Python](https://www.python.org/dev/peps/pep-0567/)). It’s also useful for defining something like logging system and reproducible [parallel PRNG](https://github.com/JuliaLang/julia/pull/34852) in “user space.”

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 3, 2020, 4:20pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/3 "2020-05-03T16:20:54Z")

</div>

> [@tkf](#):
>
> IIUC, this doesn’t work because `task_local_storage` is not propagated to sub-tasks

Yes, that’s the tricky part … one could use custom spawn/etc. constructs that do this, but I guess long term would require a change in Julia … that would make it an official standard, then. 🙂

> It’d be _awesome_ if you can come up with partr-like abstraction for distributed computation

I asked about that at JuliaCon 2019 - I think @jeff.bezanson said that something may be in the works, long term, but I don’t know any details.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 3, 2020, 4:26pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/4 "2020-05-03T16:26:04Z")

</div>

Oops, sorry, you are already aware of the problem with `task_local_storage`. Yeah, custom spawn sounds like the way to go ATM. If you want a “big gun” to solve this I guess you can also use something like IRTools.jl/Cassette.jl.

---

<div class="post-metadata">

**Author:** ![anon92994695](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@anon92994695](https://discourse.julialang.org/u/anon92994695)\
**Post date:** [May 3, 2020, 5:00pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/5 "2020-05-03T17:00:21Z")

</div>

Soooo I tried kicking off a project to do pretty much this… I ended up quitting on it because I didn’t have the time to follow it through. It’s basically a scheduling problem. Doubtful but maybe you can see how I tried doing this sorta thing and glean some insights or something, maybe not though

> **[MrPhelps/MrPhelps at master · caseykneale/MrPhelps](https://github.com/caseykneale/MrPhelps/tree/master/MrPhelps)**
>
> master/MrPhelps

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 3, 2020, 5:30pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/6 "2020-05-03T17:30:07Z")

</div>

> [@tkf](#):
>
> If you want a “big gun” to solve

Uhm, yes … let’s hope we can avoid to go for the heavy artillery. 🙂

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 6:39pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/7 "2020-05-05T18:39:05Z")

</div>

Actually, it wasn’t super hard to use IRTools to implement context variables

> <https://github.com/tkf/ContextVariablesX.jl/blob/623f83c40511eeb4a9c7446913d92c783f75bcf5/test/runtests.jl#L5-L17>

Current API/implementation is horrible as it copies `Dict` all the time. I think I need to look at PEP 567 for some ideas. Maybe using HAMT like they do.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 6:57pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/8 "2020-05-05T18:57:28Z")

</div>

I guess using Cassette could make propagation of context information seamless - but I’m worried about a high compile-time cost when using this on a largish code base.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 7:07pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/9 "2020-05-05T19:07:50Z")

</div>

Yeah, that’s true.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 7:25pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/10 "2020-05-05T19:25:46Z")

</div>

Can a task get it’s parent task? Maybe we can look up worker assignments recursively, up the task hierarchy, without propagating explicitly (except when they change)?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 7:37pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/11 "2020-05-05T19:37:06Z")

</div>

Parent task can be in any thread. So accessing its `task_local_storage` is a data race.

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [May 5, 2020, 7:49pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/12 "2020-05-05T19:49:11Z")

</div>

Talking about schedulers, this new project called Flux may help 🙂  
Cpuld this be what you are looking for?

In Flux, each job is a complete instance of the framework, meaning the individual task can support parallel tools, monitoring, and even launch sub-jobs that are, like fractals, smaller images of the parent job.

[Flux | Computing](https://computing.llnl.gov/projects/flux-building-framework-resource-management)

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 7:55pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/13 "2020-05-05T19:55:31Z")

</div>

> [@tkf](#):
>
> Parent task can be in any thread. So accessing its `task_local_storage` is a data race.

Oh, that I had planned to handle by keeping a global dict with task-ids and resources, with a lock to prevent race conditions. But how do I get the parent of a task?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 8:13pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/14 "2020-05-05T20:13:55Z")

</div>

Ah, OK. I didn’t realize that. Yeah it sounds like it’d work if we had access to the parent task. But, no, I don’t think there is a way to access parent task. I don’t think it’s a good idea in general as you’d want to GC the task as soon as it finishes.

Also, why a global lock? Propagating immutable persistent data structure seems to be a much better option for this to me (e.g., for reducing the contention).

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 8:16pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/15 "2020-05-05T20:16:05Z")

</div>

Let’s see if building task-context propagation into Julia itself could be acceptable: [Propagating context information to child-tasks and remote calls? · Issue #35757 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/35757)

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 8:18pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/16 "2020-05-05T20:18:26Z")

</div>

Thanks for opening an issue!

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 8:18pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/17 "2020-05-05T20:18:52Z")

</div>

> [@tkf](#):
>
> Also, why a global lock? Propagating immutable persistent data structure seems to be a much better option

Indeed - but I was thinking that there might need to be a central place to look up resource allocation, but mutable since tasks pop up and close all the time. The information would be immutable for each task, of course. In any case, I assume quite a few locks are locked and unlocked somewhere every time a task is spawned … but it was just a rough idea, not a design concept. 🙂

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [May 5, 2020, 8:20pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/18 "2020-05-05T20:20:33Z")

</div>

> [@tkf](#):
>
> Thanks for opening an issue!

Thanks!

I guess the response to that would determine a bit how to proceed (temporary workaround vs. long-term independent solution).

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 5, 2020, 8:50pm UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/19 "2020-05-05T20:50:47Z")

</div>

> [@johnh](#):
>
> this new project called Flux may help

Is Flux (very confusing name in Julia world 🙂 ) something like Slurm but with more programmable API?

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [May 6, 2020, 9:31am UTC](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670/20 "2020-05-06T09:31:19Z")

</div>

I believe so. Flux is quite new. It is intended to implement more fine grained control of how jobs are run.  
So yes I think it will fit in very well with the use case here.

Also if you use Slurm at the moment, Flux will run ‘on top’ of Slurm.

[Next page](https://discourse.julialang.org/t/propagation-of-available-assigned-worker-ids-in-hierarchical-computations/38670.md?page=2)
