# @distributed retry

**URL:** https://discourse.julialang.org/t/distributed-retry/78703
**Category:** General Usage
**Tags:** question
**Created:** [March 29, 2022, 9:58pm UTC](https://discourse.julialang.org/t/distributed-retry/78703 "2022-03-29T21:58:34Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![Satvik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/satvik/32/20486_2.png) [@Satvik](https://discourse.julialang.org/u/Satvik)
#### Post date: [March 29, 2022, 9:58pm UTC](https://discourse.julialang.org/t/distributed-retry/78703/1 "2022-03-29T21:58:35Z")

</div>

I have code that looks roughly like this:

```julia
using Distributed
@everywhere using MyModule
@sync @distributed for x in xs
    my_function!(x)
end

```

Which I launch using `julia -p 4 --threads=auto myscript.jl`. `my_function` writes a file to disk and doesn’t return anything.

`my_function` sometimes crashes, due to hard-to-predict memory allocation or other issues. When that happens, I’d prefer to retry it on the same input where it crashed. The problem is that the crash usually terminates the worker, e.g.

```julia
Worker 4 terminated.
Unhandled Task ERROR: EOFError: read end of file

```

So putting retry logic in `my_function` won’t help. Is there a straightforward way to e.g. assign the failed value to another worker and retry?

---

<div class="post-metadata">

### Author: ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)
#### Post date: [March 29, 2022, 10:04pm UTC](https://discourse.julialang.org/t/distributed-retry/78703/2 "2022-03-29T22:04:33Z")

</div>

To be honest, I’ve seen better MWE’s;) Is it the same file all workers write (that would explain a lot)?

> [@Satvik](#):
>
> `my_function` sometimes crashes, due to hard-to-predict memory allocation or other issues.

Maybe you should analyze these issues first? In any case you could alleviate these problems with `try`-`catch`-handling in the worker?

---

<div class="post-metadata">

### Author: ![Satvik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/satvik/32/20486_2.png) [@Satvik](https://discourse.julialang.org/u/Satvik)
#### Post date: [March 29, 2022, 10:16pm UTC](https://discourse.julialang.org/t/distributed-retry/78703/3 "2022-03-29T22:16:48Z")

</div>

> [@goerch](#):
>
> Is it the same file all workers write (that would explain a lot)?

No, they all output different files.

> [@goerch](#):
>
> Maybe you should analyze these issues first? In any case you could alleviate these problems with `try` - `catch` -handling in the worker?

The real situation, of course, is more complex – there are dozens of workers each performing long-running jobs, and occasionally when all of them want a lot of memory one will crash. Running a single worker version never crashes, and while I’ve spent some time optimizing memory usage already, I’ve never seen a distributed system in real life that doesn’t have _some_ spurious crashes, and require retrying.

As an example, `Distributed.pmap` has retrying built in, so I was wondering if there’s a nice way to do something similar with the `@distributed` macro.

---

<div class="post-metadata">

### Author: ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)
#### Post date: [March 29, 2022, 10:28pm UTC](https://discourse.julialang.org/t/distributed-retry/78703/4 "2022-03-29T22:28:25Z")

</div>

> [@Satvik](#):
>
> The real situation, of course, is more complex – there are dozens of workers each performing long-running jobs, and occasionally when all of them want a lot of memory one will crash.

OK, this sounds like a design problem.

> [@Satvik](#):
>
> Running a single worker version never crashes, and while I’ve spent some time optimizing memory usage already, I’ve never seen a distributed system in real life that doesn’t have _some_ spurious crashes, and require retrying.

This sounds reasonable: in a large enough cluster some node could randomly fail.

> [@Satvik](#):
>
> As an example, `Distributed.pmap` has retrying built in…

OK. here is the [source](https://github.com/JuliaLang/julia/blob/62e0729dbc5f9d5d93d14dcd49457f02a0c6d3a7/stdlib/Distributed/src/macros.jl#L330-L361). `@distributed` seems to be using [`pfor`](https://github.com/JuliaLang/julia/blob/62e0729dbc5f9d5d93d14dcd49457f02a0c6d3a7/stdlib/Distributed/src/macros.jl#L330-L361) which has some `error_monitor`, but seemingly not the same retry logic like `pmap`.
