# Configure Slurm File for Parallelization

**URL:** <https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239>\
**Category:** Julia at Scale\
**Tags:** parallel, cluster, distributed, slurm\
**Created:** [October 13, 2024, 2:13am UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239 "2024-10-13T02:13:42Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![benjamin.nunez](https://avatars.discourse-cdn.com/v4/letter/b/74df32/32.png) [@benjamin.nunez](https://discourse.julialang.org/u/benjamin.nunez)\
**Post date:** [October 13, 2024, 2:13am UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239/1 "2024-10-13T02:13:42Z")

</div>

**Hello everyone!**

I am using a computer cluster to parallelize the execution of my Julia code.

The partition I am using has a total of 14 nodes, and my goal is to connect to this partition and execute my `main.jl` file, which specifies the following lines at the beginning:

```julia
using Distributed
using ClusterManagers
addprocs(SlurmManager(14))

```

My question is: how should I configure the Slurm script to ensure the code parallelizes correctly? Currently, my Slurm script contains the following lines:

```bash
#SBATCH --tasks=1
#SBATCH --nodes=14

```

I understand that the number of tasks should be 1, as I am only submitting the job to execute my `main.jl` file, but setting `--nodes=14` generates an error indicating that more processors were requested than available. I believe there might be a configuration with `#SBATCH --cpus-per-task=M`, but I am unsure if this applies in my case.

I would appreciate any advice.

---

<div class="post-metadata">

**Author:** ![hexaeder](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexaeder/32/24403_2.png) [@hexaeder](https://discourse.julialang.org/u/hexaeder)\
**Post date:** [October 13, 2024, 3:39pm UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239/2 "2024-10-13T15:39:46Z")

</div>

This doesn’t necessarily answer your question, but consider checking out [SlurmClusterManager.jl](https://github.com/kleinhenz/SlurmClusterManager.jl) , which leaves the resource allocation entirely to slurm and the `addprocs` inside Julia without additional arguments will just add procs according to your slurm allocation. I found it much more straightforward that way…

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [October 13, 2024, 4:33pm UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239/3 "2024-10-13T16:33:42Z")

</div>

You do not need a batch file if you are using `ClusterManagers`. When you run `addprocs(SlurmManager(14))` it runs an `srun` command internally to allocate available resources. If you need more fine-control over the allocation (like `--nodes`), you can pass in more arguments to `SlurmManager`, see documentation. Once you do `addprocs(SlurmManager(14))`, you can verify Slurm has allocated resources by running `sinfo` or `squeue` in the terminal.

From Julia’s point of view, you now have `N = 14` distributed workers. The easiest way to parallelize your code is to use `pmap`. You code can look something like the following:

```julia
function long_running_simulation(simid) 
   sleep(5) 
   println("running $simid on host $gethostname()")
   return value
end

function run_simulation(simid) 
results = pmap(1:n_sims) do x 
   long_running_simulation(x) 
end 
end 

```

This will run independent copies of `long_running_simulations` over processors as they are available. So for example, if `n_sims = 100` then it will initially run 14 `long_running_simulations`, and as each simulation finishes and a processor becomes available, it will run `long_running_simulations()` again, and so on. The variable `results` is an array of the the return values of `long_running_simulations`.

Note: There is some additional work to be done like running `@everywhere` to make sure `long_running_simulations` is actually available on the worker processes.

---

<div class="post-metadata">

**Author:** ![benjamin.nunez](https://avatars.discourse-cdn.com/v4/letter/b/74df32/32.png) [@benjamin.nunez](https://discourse.julialang.org/u/benjamin.nunez)\
**Post date:** [October 13, 2024, 7:35pm UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239/4 "2024-10-13T19:35:37Z")

</div>

Thanks @affans for your answer. I had no idea that when using ClusterManagers.jl, it’s not necessary to manually create a “.slurm” file. On the other hand, my parallelization specifically involves using the `pmap()` function and `@everywhere` to define my parallelizable function on all workers.

However, I still have one question, and I would appreciate it if you could help me.

Since I’m not creating a slurm file, arguments like the partition name, the maximum execution time, etc., I understand that they are passed as kwargs to `addprocs(SlurmManager(14))`. Do you know where I can find a list of possible kwargs that this function accepts?

Once again, thank you very much!

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [October 13, 2024, 7:41pm UTC](https://discourse.julialang.org/t/configure-slurm-file-for-parallelization/121239/5 "2024-10-13T19:41:57Z")

</div>

You can pass any `srun` [supported options](https://slurm.schedmd.com/srun.html). In the [source code](https://github.com/JuliaParallel/ClusterManagers.jl/blob/master/src/slurm.jl), the string

```julia
srun_cmd = `srun -J $jobname -n $np -D $exehome $(srunargs) $exename $exeflags $(worker_arg())

```

is built programatically where `srunargs` are what you passed in. So for example, you can have something like:

```julia
addprocs(SlurmManager(2), partition="debug", t="00:5:00")

```

which uses that particular partition. You can either use shorthand options (i.e., `t = ...` or longhand options like `partition = ...`).
