# Error when using @distributed for on cluster with multiple nodes

**URL:** <https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377>\
**Category:** Julia at Scale\
**Tags:** cluster\
**Created:** [August 31, 2018, 5:58pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377 "2018-08-31T17:58:24Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![thehalfspace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thehalfspace/32/7217_2.png) [@thehalfspace](https://discourse.julialang.org/u/thehalfspace)\
**Post date:** [August 31, 2018, 5:58pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377/1 "2018-08-31T17:58:24Z")

</div>

I have a code with some simple `@sync @distributed for` loops. It works fine when I run it on my computer with 4 processors, or on the cluster with 1 node and any number of processors per node.

But when I run it with more than one node, it gives me error:

```julia
ERROR: LoadError: On worker 7:
LoadError: peer 8 didn't connect to 7 within 59.99999... seconds

```

I am only using SharedArray and @distributed for loops. Any suggestions?

---

<div class="post-metadata">

**Author:** ![Pbellive](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pbellive/32/3604_2.png) [@Pbellive](https://discourse.julialang.org/u/Pbellive)\
**Post date:** [August 31, 2018, 6:38pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377/2 "2018-08-31T18:38:02Z")

</div>

It’s hard to say what the issue is without more information. Have you checked that you have successfully launched Julia worker processes on multiple nodes? Just a guess but one potential culprit might be ssh tunneling.

Which version of Julia are you using? Haven’t tried out 1.0 in a multiple machine setting yet but on 0.64 on my research’s group’s cluster, I fail to connect to workers on nodes other than the one hosting the master process using `addprocs` if I don’t indicate that ssh tunneling is required. E.g. (where tera31 and tera32 are hostnames of two nodes) for me

```julia
procs = ["tera31","tera32"]
addprocs(procs)

```

fails but

```julia
addprocs(procs,tunnel=true)

```

works. If you’re in an environment where it takes a long time for the connections with remote workers to be established for whatever reason, you can also try setting the `JULIA_WORKER_TIMEOUT` environment variable on the master process before calling addprocs. This will make Julia wait longer before giving up on connecting to workers.

---

<div class="post-metadata">

**Author:** ![thehalfspace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thehalfspace/32/7217_2.png) [@thehalfspace](https://discourse.julialang.org/u/thehalfspace)\
**Post date:** [August 31, 2018, 6:48pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377/3 "2018-08-31T18:48:29Z")

</div>

@Pbellive I was just doing

```julia
#PBS -l pmem=8gb,nodes=4:ppn=4,walltime=20:00:100 
julia -p 16 run.jl

```

on my PBS script. I am not adding procs later.

I am using julia1.0.0 but I think you have correctly pointed out my mistake. How do I get the hostnames? I thought the nodes are assigned randomly to my job depending on the availability.

---

<div class="post-metadata">

**Author:** ![Pbellive](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pbellive/32/3604_2.png) [@Pbellive](https://discourse.julialang.org/u/Pbellive)\
**Post date:** [August 31, 2018, 8:23pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377/4 "2018-08-31T20:23:31Z")

</div>

Oh, yes, if you’re using a job management system you’ll have to manage things a bit differently. I’m not familiar with PBS. You might want to checkout the thread:

> [@Help setting up Julia on a cluster](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519):
>
> I am trying Julia for the first time on a HPC cluster at my university. Julia v0.6 is already installed, I have a few questions: How to install packages on a custom folder in the cluster other than .julia/v0.6, and without internet connection? The cluster is configured with PBS for resource management. I found the [ClusterManagers.jl](https://github.com/JuliaParallel/ClusterManagers.jl) package, but I wonder what is the workflow? Do I still need to write a PBS script and call Julia from there?

and also the [ClusterManagers](https://github.com/JuliaParallel/ClusterManagers.jl) package.

I can say that just launching julia via `julia -p n` for some number `n`. is meant for launching multiple workers on a single machine. To launch workers on multiple machines you need to launch julia with a machine file. [This post](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/6) has an example of how to do that with PBS. That’s about all I know. If that doesn’t get you going I would look around for more resources on/ask for help with getting Julia working with cluster job schedulers.

Cheers, Patrick

---

<div class="post-metadata">

**Author:** ![thehalfspace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thehalfspace/32/7217_2.png) [@thehalfspace](https://discourse.julialang.org/u/thehalfspace)\
**Post date:** [August 31, 2018, 9:02pm UTC](https://discourse.julialang.org/t/error-when-using-distributed-for-on-cluster-with-multiple-nodes/14377/5 "2018-08-31T21:02:24Z")

</div>

awesome, thanks! I”ll check that out.
