# JULIA\_WORKER\_TIMEOUT error

**URL:** <https://discourse.julialang.org/t/julia-worker-timeout-error/24866>\
**Category:** Julia at Scale\
**Created:** [June 2, 2019, 9:59pm UTC](https://discourse.julialang.org/t/julia-worker-timeout-error/24866 "2019-06-02T21:59:13Z")\
**Posts on this page:** 1\
**Showing post:** 1

<div class="post-metadata">

**Author:** ![abx](https://avatars.discourse-cdn.com/v4/letter/a/edb3f5/32.png) [@abx](https://discourse.julialang.org/u/abx)\
**Post date:** [June 2, 2019, 9:59pm UTC](https://discourse.julialang.org/t/julia-worker-timeout-error/24866/1 "2019-06-02T21:59:13Z")

</div>

I am working on a fairly big code base written in Julia 1.1 and I am experiencing erratic issues with parallel processing. It is very annoying as most of the time everything works perfectly.  
The application is mainly composed by a long list of tasks distributed with pmap, some of them make use of distributed aggregation similar to this

```julia
result = @distributed (+) for task_id in task_ids
    aggregate_worker(task_id)
end

```

just once every many calculation I get this error from the processes that make use the aggregation code above:

_peer 36 didn’t connect to 27 within 180.0 seconds_  
_peer 36 didn’t connect to 28 within 180.0 seconds_  
_peer 36 didn’t connect to 13 within 180.0 seconds_  
_peer 36 didn’t connect to 29 within 180.0 seconds_

…and so on. You might have noticed that I have already extended the JULIA\_WORKER\_TIMEOUT to accomodate longer executions times. That did not help much. All the processes run on the same machine.

If I resubmit the failing task, that gets processed correctly. I know it is a long shot with the little information given, however any idea/recommendation on how to investigate or where to look at would be very appreciated. Thanks!

---

_[View the full topic](https://discourse.julialang.org/t/julia-worker-timeout-error/24866)._
