# Help setting up Julia on a cluster

**URL:** <https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519>\
**Category:** Julia at Scale\
**Tags:** question, parallel, cluster\
**Created:** [August 23, 2017, 4:22am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519 "2017-08-23T04:22:37Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 23, 2017, 4:22am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/1 "2017-08-23T04:22:37Z")

</div>

I am trying Julia for the first time on a HPC cluster at my university. Julia v0.6 is already installed, I have a few questions:

1. How to install packages on a custom folder in the cluster other than `.julia/v0.6`, and without internet connection?
2. The cluster is configured with PBS for resource management. I found the [ClusterManagers.jl](https://github.com/JuliaParallel/ClusterManagers.jl) package, but I wonder what is the workflow? Do I still need to write a PBS script and call Julia from there?

---

<div class="post-metadata">

**Author:** ![ararslan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ararslan/32/3825_2.png) [@ararslan](https://discourse.julialang.org/u/ararslan)\
**Post date:** [August 23, 2017, 5:24am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/2 "2017-08-23T05:24:15Z")

</div>

To address 1, assuming you have a networked file system, you should be able to put your packages in a directory accessible by every node then add

```julia
prepend!(LOAD_PATH, "/some/path")

```

in a .juliarc.jl file on the head node.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 23, 2017, 6:03am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/3 "2017-08-23T06:03:43Z")

</div>

[https://docs.julialang.org/en/latest/manual/packages/#Offline-Installation-of-Packages-1](https://docs.julialang.org/en/latest/manual/packages/#Offline-Installation-of-Packages-1)

---

<div class="post-metadata">

**Author:** ![John\_Hearns](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/john_hearns/32/1685_2.png) [@John\_Hearns](https://discourse.julialang.org/u/John_Hearns)\
**Post date:** [August 23, 2017, 2:37pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/4 "2017-08-23T14:37:45Z")

</div>

@juliohm I asked a similar question recently. You might find this useful:

> [@Package management - an intro?](https://discourse.julialang.org/t/package-management-an-intro/4344):
>
> I see that there is a talk by Stefan Karpinski on Pkg3 package management. Python ‘virtualenv’ is mentioned. Would anyone be kind enough to point me towards a gentle intro to Julia package management. My specific interest is managing Julia installaitons and package on HPC clusters, where there is traditionally a central installation on a shared network drive. However people can and do want to install specific packages or bleeding edge versions for their own use.

---

<div class="post-metadata">

**Author:** ![John\_Hearns](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/john_hearns/32/1685_2.png) [@John\_Hearns](https://discourse.julialang.org/u/John_Hearns)\
**Post date:** [August 23, 2017, 3:14pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/5 "2017-08-23T15:14:00Z")

</div>

@juliohm WRT your second question, an expert will be along in a minute.  
The examples from ClusterManagers.jl have the paradigm that you start a Julia session, then use ClusterManagers to submit a job (or start processes locally using affinity). That job will be started with the number of cores you request, int he queue you request and with the other resources you request in the qsub\_env  
I think I may well have this totally wrong, but if you want this to run totally in a batch mode you would either have to start the main script with the master process in a terminal using a screen session and leave it running.  
Or you could qsub the main/master script which in turn submits the ClusterManager job.

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [August 23, 2017, 7:40pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/6 "2017-08-23T19:40:19Z")

</div>

I recently went through the same learning process on my university’s cluster, and your second question was the hardest part for me. I found the easiest way was to submit to one of the queues using a PBS submit script that tells Julia where to find all the processors via the PBS nodefile. Here’s a minimum working example PBS script:

```julia
#!/bin/bash
#PBS -l nodes=4:ppn=12,walltime=00:05:00
#PBS -N test_julia
#PBS -q debug

echo PBS: node file is $PBS_NODEFILE
julia --machinefile=$PBS_NODEFILE /path/to/your/home/dir/test_julia.jl
echo "finished"

```

I called this file `test_julia.pbs`. When you submit this job (i.e., by running `qsub test_julia.pbs` from your login prompt) the Julia process starts up with all the processors (48 in this case) available, as if you’d started it on your laptop with `julia -p 2` or whatever. For completeness, here’s `test_julia.jl`, which has minimum working examples for basic batch-processing tasks:

```julia
println("Hello from Julia")
np = nprocs()
println("Number of processes: $np")

for i in workers()
    host, pid = fetch(@spawnat i (gethostname(), getpid()))
    println("Hello from process $(pid) on host $(host)!")
end

tasks = randn(np * 30)

@everywhere begin
    function foo(x)
        return x * 4
    end
end

results = pmap(foo, tasks)

println(round(results, 3))

for i in workers()
    rmprocs(i)
end

```

Hope that helps!

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 23, 2017, 7:49pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/7 "2017-08-23T19:49:57Z")

</div>

Thank you @ElOceanografo, that is really helpful already. I can give it a try and get things moving while I digest the workflow, possibly with ClusterManagers.jl.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 25, 2017, 9:23pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/8 "2017-08-25T21:23:32Z")

</div>

@ElOceanografo, I am getting an error saying that the file /home/juliohm/.ssh cannot be created. This makes sense because in the cluster I am trying to run, users cannot write to /home. Did anyone had a similar issue?

Also, for requesting more than one node on clusters, do we need to install MPI.jl? Is MPI.jl deprecated by ClusterManagers.jl?

If MPI.jl is required, that is bad news. It only supports Julia v0.4 according to the README.

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [August 25, 2017, 10:16pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/9 "2017-08-25T22:16:21Z")

</div>

The file permission error is something I would ask to your cluster’s admin about. On the cluster I’m using, we’re given write access in out home directory, so I didn’t run into this problem.

As to your second question, MPI.jl is not required. You don’t need to use ClusterManagers either if you don’t want to. As I understand it, when you run `addprocs_pbs()` from that package, it just [auto-generates the appropriate PBS commands](https://github.com/JuliaParallel/ClusterManagers.jl/blob/master/src/qsub.jl#L71) within Julia and sends them off via `qsub`. I tried this approach and couldn’t get it to work…`addprocs_pbs` would hang. I didn’t spend much time investigating why, because the slightly-less-elegant solution I posted above, with a separate submit script, worked fine.

I’m not actually an expert in any of this, though, so if other, more-informed people want to chime in that would be great.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 25, 2017, 10:29pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/10 "2017-08-25T22:29:25Z")

</div>

I asked the cluster’s admin and he mentioned that the cluster doesn’t use SSH for communication, which I think is bad news for me. Julia will only be able to use one node in this case. Please correct me if I am wrong.

He also mentioned that a MPI manager is required to have processes running across different nodes, is there any documentation on how these concepts are connected?

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [August 25, 2017, 11:15pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/11 "2017-08-25T23:15:12Z")

</div>

Ah…you may have to use MPI after all, in that case. I would see if you can get any of the MPI.jl examples [here](https://github.com/JuliaParallel/MPI.jl/tree/master/examples) to run–you’ll use a `qsub` submit script that looks something like this:

```julia
#!/bin/bash
#PBS -l nodes=4:ppn=12,walltime=00:05:00
#PBS -N test_julia
#PBS -q debug

module load openmpi
mpirun -np 48 /path/to/your/home/dir/01-hello.jl
echo "finished"

```

You may need to change `module load openmpi` to refer to whatever version of MPI is available on your cluster (or you may be able to omit it if MPI is loaded by default). I don’t _know_ if this will work, but I have used a similar pattern to successfully run Python programs using `mpi4py`, the package which inspired MPI.jl.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 26, 2017, 1:01am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/12 "2017-08-26T01:01:02Z")

</div>

Yes, I used mpi4py in the past, I assumed that most people doing HPC in Julia were using it also, but apparently all these threads refer to SSH communication? What about the Celeste project, do you know if they used MPI?

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [August 26, 2017, 7:01pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/13 "2017-08-26T19:01:33Z")

</div>

It looks like they used a new-to-me library called [Gasp.jl](https://github.com/kpamnany/Gasp.jl), which wraps a C library, also called Gasp, which [does rely on MPI](https://github.com/kpamnany/gasp/blob/master/src/gasp.c#L7).

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 26, 2017, 10:46pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/14 "2017-08-26T22:46:43Z")

</div>

It looks like I misinterpreted the type of parallel computing that Julia supports natively. By just reading the [documentation](https://docs.julialang.org/en/latest/manual/parallel-computing), it compares Julia’s message passing with MPI message passing, which is not a fair comparison at all.

MPI is about distributed-memory systems where nodes communicate through high-performance hardware like [InfiniBand](https://en.wikipedia.org/wiki/InfiniBand) techonology. As far as I can tell, Julia’s parallel computing features do not support any of that.

So all the beautiful `pmap`, `@parallel` and related functions are useless in an HPC cluster with hundreds of nodes? I think we need a generalization of parallel programming paradigms in a package that implements the `pmap` operation on arbitrary pools of processes. I will try to work on this if I find the time, `pmap` covers more than 80% of use cases out there for embarrassingly parallel tasks, and it looks like there is no clean solution yet.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [August 27, 2017, 7:31am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/15 "2017-08-27T07:31:45Z")

</div>

> [@juliohm](#):
>
> MPI is about distributed-memory systems where nodes communicate through high-performance hardware like InfiniBand techonology. As far as I can tell, Julia’s parallel computing features do not support any of that.

Yes it does. It is distributed memory across multiple nodes. You can configure in more detail how it communicates through the ClusterManager I think.  
But even if you just open it up with a machine file it’ll be distributed. Here’s an example of doing it the machine file way:

> **[Multi-node Parallelism in Julia on an HPC (XSEDE Comet) - Stochastic Lifestyle](http://www.stochasticlifestyle.com/multi-node-parallelism-in-julia-on-an-hpc/)**
>
> Today I am going to show you how to parallelize your Julia code over some standard HPC interfaces. First I will go through the steps of parallelizing a simple code, and then running it with single-node parallelism and multi-node parallelism. The...

But of course you can add processes any of the other documented way and it will distribute as well. @oxinabox had a good recent blog post on a few different methods for doing this:

[http://white.ucc.asn.au/2017/08/17/starting-workers.html](http://white.ucc.asn.au/2017/08/17/starting-workers.html)

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 27, 2017, 3:29pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/16 "2017-08-27T15:29:43Z")

</div>

Hi @ChrisRackauckas, thank you for the links, but what about the SSH issue? It seems that not all clusters allow nodes to communicate via SSH? I am gonna contact the admin of the cluster again to ask this in more detail.

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [August 27, 2017, 7:16pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/17 "2017-08-27T19:16:52Z")

</div>

Can confirm what @ChrisRackauckas said, since I’ve run code on multiple processors using `pmap` myself. It does sound like the SSH restriction may be the issue on your cluster…so following up with the admin is probably the best next step.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [October 27, 2017, 1:21am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/18 "2017-10-27T01:21:25Z")

</div>

I want to add an update and solution to this thread.

The answer by @ElOceanografo is great, specially if you are interested in fine control with the cluster resource manager. I would say that the most clean solution nowadays is through the ClusterManagers.jl package.

Basically, the package defines various functions `addprocs_slurm()`, `addprocs_pbs()`, etc. which can be used in place of the Julia builtin `addprocs()`. When you start Julia, instead of asking for processes with `addprocs()`, you simply do:

```julia
using ClusterManagers

addprocs_pbs(100) # request 100 processes in the cluster using PBS

# REST OF SCRIPT GOES HERE

```

This will take care of submitting the job with the appropriate resource manager without having to write the job script manually.

---

<div class="post-metadata">

**Author:** ![Anna\_W](https://avatars.discourse-cdn.com/v4/letter/a/e5b9ba/32.png) [@Anna\_W](https://discourse.julialang.org/u/Anna_W)\
**Post date:** [January 23, 2018, 10:33am UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/19 "2018-01-23T10:33:08Z")

</div>

Hello, Does anyone has a expirience with running Julia on cluster under the Torque schedguler. The Cluster manager with addprocs\_pbs() do not work.

---

<div class="post-metadata">

**Author:** ![thehalfspace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thehalfspace/32/7217_2.png) [@thehalfspace](https://discourse.julialang.org/u/thehalfspace)\
**Post date:** [September 2, 2018, 4:20pm UTC](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/20 "2018-09-02T16:20:32Z")

</div>

Have you tried this in julia 1.0.0? I am unable to figure it out, though the docs have something about [cluster manager interface](https://docs.julialang.org/en/v1/stdlib/Distributed/#Distributed.manage). I can’t find examples on how to implement it. Any ideas?

[Next page](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519.md?page=2)
