# Understanding data movement with Distributed.jl

**URL:** https://discourse.julialang.org/t/understanding-data-movement-with-distributed-jl/100680
**Category:** General Usage
**Tags:** distributed
**Created:** [June 22, 2023, 12:55am UTC](https://discourse.julialang.org/t/understanding-data-movement-with-distributed-jl/100680 "2023-06-22T00:55:22Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![lmtzx9h4qqnt](https://avatars.discourse-cdn.com/v4/letter/l/74df32/32.png) [@lmtzx9h4qqnt](https://discourse.julialang.org/u/lmtzx9h4qqnt)
#### Post date: [June 22, 2023, 12:55am UTC](https://discourse.julialang.org/t/understanding-data-movement-with-distributed-jl/100680/1 "2023-06-22T00:55:22Z")

</div>

As outlined in the manual, [Multi-processing and Distributed Computing · The Julia Language](https://docs.julialang.org/en/v1/manual/distributed-computing/#Data-Movement), Distributed.jl performs “implicit” data movement operations between the different processes. Unfortunately, that page doesn’t really make it clear under what conditions these occur and how to have some control over the communication between processes.

Generally, I think three different types of variables are relevant here: global variables, arguments passed to remotely-called functions, and closure variables.

- **Global variables:** This seems like the most difficult one to follow (it doesn’t help that, for instance, the manual incorrectly refers to a function referring to a global variable as a “closure”). This seems like the most pertinent part:

- **Arguments passed to remotely-called functions:** I assume these are always just serialized, sent to the remote process, and passed to the function there (without defining any globals)?

- **Closures:** Any closed-over variables are presumably included in the serialization process and thus sent just like function arguments?

So, do I generally have the right idea here? Any answers to the sub-questions above would be appreciated as well!

---

<div class="post-metadata">

### Author: ![lmtzx9h4qqnt](https://avatars.discourse-cdn.com/v4/letter/l/74df32/32.png) [@lmtzx9h4qqnt](https://discourse.julialang.org/u/lmtzx9h4qqnt)
#### Post date: [September 2, 2024, 11:00pm UTC](https://discourse.julialang.org/t/understanding-data-movement-with-distributed-jl/100680/2 "2024-09-02T23:00:02Z")

</div>

Regarding the question of why “named” functions behave differently from anonymous functions for serialization, I just came across this post in a different thread:

> [@\`foo = () -\> 3\` vs. \`foo() = 3\`](https://discourse.julialang.org/t/foo-3-vs-foo-3/59710/8):
>
> Just for completeness, another small difference is how they’re serialized. For the anonymous version, the function definition is actually serialized, where as the non-anonymous one just the name is serialized. The niche place where this might come up is when using Distributed’s auto-global shipping, so e.g. this works, foo = () -\> 3 @fetch foo() # foo definition shipped to worker, run there, answer returned whereas the non-anonymous version you need to have foo actually defined on all workers …
