# How to debug remotecall\_wait/deserialization failures?

**URL:** <https://discourse.julialang.org/t/how-to-debug-remotecall-wait-deserialization-failures/77864>\
**Category:** Julia at Scale\
**Tags:** question, distributed, tcp\
**Created:** [March 14, 2022, 1:18pm UTC](https://discourse.julialang.org/t/how-to-debug-remotecall-wait-deserialization-failures/77864 "2022-03-14T13:18:54Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![jonas-schulze](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jonas-schulze/32/9228_2.png) [@jonas-schulze](https://discourse.julialang.org/u/jonas-schulze)\
**Post date:** [March 15, 2022, 2:54pm UTC](https://discourse.julialang.org/t/how-to-debug-remotecall-wait-deserialization-failures/77864/2 "2022-03-15T14:54:21Z")

</div>

After reading a bit of [Failure-resilient parallel computing](https://discourse.julialang.org/t/failure-resilient-parallel-computing/17929) and the documentation of `pmap`, I wrapped by code in `retry`:

```julia
tasks = map(workers) do w
    @async retry(
        () -> remotecall_wait(w, ...);
        relays = ExponentialBackOff(n=3),
        check = #= that the error is not mine =#,
    )()
end

```

On the workers, everything is already wrapped in a `try catch` block. That makes the `check` part of `retry` a bit easier. Overall, I get over the first error described in the original post. Some stages are re-tried twice, so `n=3` might not even be enough.

However, when writing the data back to disk (using [`HDF5.jl`](https://github.com/JuliaIO/HDF5.jl) v0.15.7 and Julia v1.6.1, writing to one file per process), I almost allways hit the Slurm timeout. The logs of stdout/stderr are then full of

```julia
...
signal (15): Terminated
in expression starting at none:0
epoll_wait at /lib64/libc.so.6 (unknown line)

signal (15): Terminated
in expression starting at none:0
...

```

without any of my own prints/logs. The other logfiles that I write (one per process) are ok and show that my algorithm executed properly. As I don’t handle errors that thoroughly on the “store my results” part, I suspected that maybe some worker process died for a failed `remotecall_fetch` (similar to the issue in the original post). So I issued the job once more but with a stupidly big timeout to check whether the workers are still alive after the actual algorithm was done. Unfortunately, they are:

```bash
$ cat dead-d44.nodes
node043
node044
node045
...

$ cat dead-d44.nodes | xargs -P0 -I{} -n1 ssh {} 'ps -U `whoami` -ocmd= | ...grep for workers... | wc -l' > dead-d44.counts

$ awk '{sum+=$1;} END{print sum;}' dead-d44.counts
450

```

(which I should have guessed, because Slurm caught the failures and aborted the whole job in the original issue … it did not this time)

The `epoll_wait` hints at `HDF5.jl` or maybe `libuv`? Julia v1.6.1 is pretty dated, so I am trying v1.6.5 now. What else can I do?

---

_[View the full topic](https://discourse.julialang.org/t/how-to-debug-remotecall-wait-deserialization-failures/77864)._
