# Issues with machinefile and SLURM

**URL:** <https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882>\
**Category:** Julia at Scale\
**Created:** [December 20, 2017, 4:47pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882 "2017-12-20T16:47:55Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![nickeubank](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nickeubank/32/3878_2.png) [@nickeubank](https://discourse.julialang.org/u/nickeubank)\
**Post date:** [December 20, 2017, 4:47pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/1 "2017-12-20T16:47:55Z")

</div>

Spent last week fighting with issues on a slurm cluster, and having finally figured it out, I wanted to share the result (and a warning):

I had been parallelizing through a slurm batch script with this call:

`julia --machinefile $SLURM_NODEFILE indiv_array.jl`

(For full script, [see here](https://discourse.julialang.org/t/debugging-possible-issue-with-machinefile-option-on-slurm-system/7857))

Turns out (The Vanderbilt IT team and I have discovered) the problem with this strategy is that because julia opens new processes using `ssh`, they escape SLURM’s notice. As a result, my parallel workers were running outside of SLURMs awareness (technical term is, I believe, outside the `cgroup`), taking up memory unexpectedly and not always shutting down when `scancel` was called on the main task (at one point I apparently had \>30 zombie processes running on the research cluster, even though my `squeue` was clean).

I think `ClusterManager.jl` solves this, but doesn’t seem to work well for busy slurm clusters (that generally require use of `sbatch` script and long waits for resources) since it uses `srun`.

So… oops.

CC: @ChrisRackauckas @raminammour

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 20, 2017, 7:23pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/2 "2017-12-20T19:23:33Z")

</div>

Using `srun` should be fine since that is the correct way of starting a job.

The workflow should work something like this:

```julia
salloc | sbatch # create resources.
julia> addprocs(SlurmManager(2)) # SlurmManager should inherit the outside allocation.

```

---

<div class="post-metadata">

**Author:** ![nickeubank](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nickeubank/32/3878_2.png) [@nickeubank](https://discourse.julialang.org/u/nickeubank)\
**Post date:** [December 20, 2017, 7:41pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/3 "2017-12-20T19:41:23Z")

</div>

OH! So an `srun` executed inside a slurm allocation doesn’t try to create a new allocaiton; it start processes _in that existing allocation_?

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 20, 2017, 7:53pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/4 "2017-12-20T19:53:37Z")

</div>

Yes. See [Slurm Workload Manager - srun](https://slurm.schedmd.com/srun.html)

> Run a parallel job on cluster managed by Slurm. If necessary, srun will first create a resource allocation in which to run the parallel job.

`srun` in general is the right way of starting jobs within an allocation and crucially within `sbatch`.

[https://slurm.schedmd.com/sbatch.html](https://slurm.schedmd.com/sbatch.html)

> When the job allocation is finally granted for the batch script, Slurm runs a single copy of the batch script on the first node in the set of allocated nodes.

That’s why a `sbatch` script has usually one or several `srun` command in it.

---

<div class="post-metadata">

**Author:** ![nickeubank](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nickeubank/32/3878_2.png) [@nickeubank](https://discourse.julialang.org/u/nickeubank)\
**Post date:** [December 21, 2017, 5:42pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/5 "2017-12-21T17:42:22Z")

</div>

This conversation is revelatory. Thank you!!

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [December 21, 2017, 5:46pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/6 "2017-12-21T17:46:03Z")

</div>

Huh, I didn’t know that would work either. Awesome @vchuravy. One note: does default `addprocs()` do the correct thing like `addprocs(SlurmManager(2))` when in a cluster job? What I mean is, does `addprocs()` automatically recognize that it should use the `SlurmManager` with 2 process when it’s called from a SLURM job with 2 cores, or is that asking too much?

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 21, 2017, 6:08pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/7 "2017-12-21T18:08:37Z")

</div>

That is asking to much. We would have to redefine `addprocs` when loading ClusterManager.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 21, 2017, 8:56pm UTC](https://discourse.julialang.org/t/issues-with-machinefile-and-slurm/7882/8 "2017-12-21T20:56:34Z")

</div>

A post was split to a new topic: [Issues running on a PBS cluster](https://discourse.julialang.org/t/issues-running-on-a-pbs-cluster/7900)
