# Running Julia in a SLURM Cluster

**URL:** <https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614>\
**Category:** Performance\
**Tags:** parallel, cluster, distributed\
**Created:** [September 3, 2021, 1:56am UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614 "2021-09-03T01:56:57Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Noah\_Guzman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noah_guzman/32/28831_2.png) [@Noah\_Guzman](https://discourse.julialang.org/u/Noah_Guzman)\
**Post date:** [September 3, 2021, 1:56am UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/1 "2021-09-03T01:56:57Z")

</div>

I’m looking to provide an explicit explanation of the simplest way to get a Julia program running in parallel across multiple nodes in a SLURM cluster (basically a minimal working example that illustrates the logic).

My impression so far is that there are two primary ways to run Julia in a SLURM cluster. Suppose I want to define a function and run it in a parallel for-loop on _N_ cores distributed across _M_ nodes in the cluster:

**Option 1**

The first option comes from [this Stackoverflow post](https://stackoverflow.com/questions/65250900/julia-parallel-processing-on-pbs-multiple-nodes). Basically, you can use some functions from the `ClusterManagers` package in your code and then just run Julia as normal without having to explicitly write a SLURM script.

The example program:

```julia
# File name
# slurm_example.jl
using Distributed
using ClusterManagers

# Add N workers across M nodes
addprocs_slurm(N, nodes=M, exename="/path/to/julia/bin/julia", rest of SLURM kwargs...)

# Define function
@everywhere function myFunction(args)
    Code goes here...
end

# Run function K times in parallel
@parallel for i=1:K
    myFunction(args)
end

```

As I understand it, to run this program, I would simply execute

```julia
julia slurm_example.jl

```

from the command line while logged into the cluster. Then the `addprocs_slurm` function runs the rest of the Julia code as an interactive SLURM job, the equivalent of using `srun` with the specified SLURM options.

**Option 2**

The second option, exemplified in [this post](https://discourse.julialang.org/t/help-setting-up-julia-on-a-cluster/5519/6), involves writing a SLURM script for a batch job calling Julia with the `--machinefile` flag. In this case, the example program is:

```julia
# File name
# slurm_example.jl
using Distributed
using ClusterManagers

# Define function
@everywhere function myFunction(args)
    Code goes here...
end

# Run function K times in parallel
@parallel for i=1:K
    myFunction(args)
end

# Kill the workers
for i in workers()
    rmprocs(i)
end

```

Then to run this, I would need to write and execute a separate SLURM script that looks something like this:

```julia
#!/bin/bash

#SBATCH --ntasks=N # N cores

#SBATCH --nodes=M # M nodes 

# Rest of #SBATCH flags go here...

 julia --machinefile=$SLURM_NODEFILE slurm_example.jl

```

One thing I find confusing about this example is that I don’t understand why I don’t need to do something like

```julia
 addprocs(SlurmManager(N))

```

in the Julia code? Or do I? Are there any glaring errors with this code? Is the main difference between the two options just that Option 1 is an interactive SLURM job and the other a batch job?

Thanks ahead of time for any feedback.

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [September 3, 2021, 5:24pm UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/2 "2021-09-03T17:24:20Z")

</div>

I recently set up some scripts for running Julia jobs on a Slurm cluster. All of my jobs just use a single node with multiple CPUs. I’ll describe my approach, but I’m not an expert in HPC, so I’m not sure if everything that I’m doing is 100% correct.

My approach is to write two scripts: a Slurm script and a Julia script. I currently am not using `ClusterManagers`. My mental model is that if I request one node with multiple CPUs, Slurm provisions a virtual machine with multiple cores, and Julia will be able to detect those cores automatically just like it does on my laptop. So basically all I need to do is `using Distributed; addprocs(4)` and then parallelize my code with `@distributed` or `pmap`.

# Example 1

### Slurm Script (“test\_distributed.slurm”)

```julia-auto
#!/bin/bash

#SBATCH -p <list of partition names>
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=00:05:00
#SBATCH --mail-type=ALL
#SBATCH --mail-user=<your email address>

julia test_distributed.jl

```

### Julia Script (“test\_distributed.jl”)

```julia
using Distributed

# launch worker processes
addprocs(4)

println("Number of processes: ", nprocs())
println("Number of workers: ", nworkers())

# each worker gets its id, process id and hostname
for i in workers()
    id, pid, host = fetch(@spawnat i (myid(), getpid(), gethostname()))
    println(id, " " , pid, " ", host)
end

# remove the workers
for i in workers()
    rmprocs(i)
end

```

### Output File

```julia-auto
Number of processes: 5
Number of workers: 4
2 2331013 cn1081
3 2331015 cn1081
4 2331016 cn1081
5 2331017 cn1081

```

# Example 2

In this example I run a parallel `for` loop with `@distributed`. The body of the `for` loop has a 5 minute `sleep` call. I verified that the loop iterations are in fact running in parallel by recording the run time for the whole job. The run time for this job was 00:05:19, rather than the 00:20:00 run time that would be expected if the code was running serially.

### Slurm Script (“test\_distributed2.slurm”)

```julia-auto
#!/bin/bash

#SBATCH -p <list of partition names>
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=00:30:00
#SBATCH --mail-type=ALL
#SBATCH --mail-user=<your email address>

julia test_distributed2.jl

```

### Julia Script (“test\_distributed2.jl”)

```julia
using Distributed

addprocs(4)

println("Number of processes: ", nprocs())
println("Number of workers: ", nworkers())

@sync @distributed for i in 1:4
    sleep(300)
    id, pid, host = myid(), getpid(), gethostname()
    println(id, " " , pid, " ", host)
end

for i in workers()
    rmprocs(i)
end

```

### Output File

```julia-auto
Number of processes: 5
Number of workers: 4
      From worker 2:	2 2334507 cn1081
      From worker 3:	3 2334509 cn1081
      From worker 5:	5 2334511 cn1081
      From worker 4:	4 2334510 cn1081

```

# Comments

If you’re using a `Project.toml` or `Manifest.toml` you will probably need to call `addprocs` like this:

```julia
addprocs(4; exeflags="--project")

```

I also had to jump through some hoops to run a project that had dependencies in private Github repos. I think it boiled down to instantiating a `Manifest.toml` file that contained the appropriate links to the private Github repos, but I didn’t document the full process…

---

<div class="post-metadata">

**Author:** ![Noah\_Guzman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noah_guzman/32/28831_2.png) [@Noah\_Guzman](https://discourse.julialang.org/u/Noah_Guzman)\
**Post date:** [September 3, 2021, 7:15pm UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/3 "2021-09-03T19:15:43Z")

</div>

Awesome, this is super clear. Thanks for taking the time to write it all out. I think this should work in the case of multiple nodes (if one node does not have enough cores, etc.), but even if it doesn’t it’s a step in the right direction, so it deserves to be the chosen solution.

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [September 3, 2021, 7:19pm UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/4 "2021-09-03T19:19:39Z")

</div>

Glad that was helpful. Hopefully other folks will chime in with alternative approaches. There’s definitely not a lot of good examples or tutorials on the web for using Julia on a cluster.

---

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [September 3, 2021, 8:44pm UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/5 "2021-09-03T20:44:19Z")

</div>

My preferred way is something like this:

```julia
#!/usr/bin/env sh
#SBATCH -N 10
#SBATCH -n 8
#SBATCH -o %x-%j.out
#=
srun julia $(scontrol show job $SLURM_JOBID | awk -F= '/Command=/{print $2}')
exit
# =#

using MPIClusterManagers
MPIClusterManagers.start_main_loop(MPI_TRANSPORT_ALL)

println(workers()) # should have 80 workers here across 10 nodes (controlled by -n and -N above)

```

You put this in `myscript.jl` and then `sbatch myscript.jl`.

This is using [A neat Julia/SLURM trick](https://discourse.julialang.org/t/a-neat-julia-slurm-trick/61123) and [MPIClusterManagers.jl](https://github.com/JuliaParallel/MPIClusterManagers.jl).

ClusterMangers’s `ElasticManager` is also quite useful for dynamically hooking up workers to e.g. a Jupyter session, if you prefer the interactive workflow.

---

<div class="post-metadata">

**Author:** ![SpuriousEigenstate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/spuriouseigenstate/32/208390_2.png) [@SpuriousEigenstate](https://discourse.julialang.org/u/SpuriousEigenstate)\
**Post date:** [April 10, 2024, 9:19am UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/6 "2024-04-10T09:19:01Z")

</div>

I find that doing this is easiest, and seems to work well for me:  
Putting this at the top of my `script.jl`,

```julia
#!/usr/bin/env julia

#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=00:30:00

using Distributed
...

```

and then just running `sbatch script.jl` does the job

---

<div class="post-metadata">

**Author:** ![maphdze](https://avatars.discourse-cdn.com/v4/letter/m/ea5d25/32.png) [@maphdze](https://discourse.julialang.org/u/maphdze)\
**Post date:** [April 11, 2024, 11:11am UTC](https://discourse.julialang.org/t/running-julia-in-a-slurm-cluster/67614/7 "2024-04-11T11:11:15Z")

</div>

That’s more simple! I will give a try.
