# Distributed Julia, ClusterMangers package, Slurm, and CUDA on a shared file system

**URL:** <https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663>\
**Category:** Julia at Scale\
**Tags:** gpu, cuda, distributed\
**Created:** [November 22, 2022, 7:19pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663 "2022-11-22T19:19:38Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![asaxton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/asaxton/32/22216_2.png) [@asaxton](https://discourse.julialang.org/u/asaxton)\
**Post date:** [November 22, 2022, 7:19pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663/1 "2022-11-22T19:19:38Z")

</div>

Hello,

I’m running Julia 1.7.3 on a cluster using Slurm to launch batch jobs. The cluster has a shared filesystem. I’m using the Slurm cluster manager from the ClusterManagers package to call `addproc(...)`. The node I’m running the REPL on is a login node and does NOT have a GPU, but the `procs` returned from `addproc()` DO have GPUs.

When I run `@fetechfrom procs[1] CUDA.ndevices()` it returns `1`. However, when I run `@fetechfrom procs[1] CUDA.devices()` it returns

```julia
Error showing value of type CUDA.DeviceIterator:
ERROR: CUDA error (code 100, CUDA_ERROR_NO_DEVICE)
Stacktrace:
[1] throw_api_error(res::CUDA.cudaError_enum)
   @ CUDA ~/.julia/packages/CUDA/DfvRa/lib/cudadrv/error.jl:89

```

Some other evidence that CUDA and the GPU’s are installed correctly on the remote hosts is if I run

```julia
srun --time=02:00:00 --partition=a100 --nodes=1 --cpus-per-task=16 --gpus-per-node=1 --pty /bin/bash
julia --project=MyProjDirWithCUDAPkg
julia> using CUDA
julia> CUDA.devices()
CUDA.DeviceIterator() for 1 devices:
0. NVIDIA A100 80GB PCIe

```

I’ve done this experiment and double checked that the host(s) returned from `addproc(...)` and `srun bash` are the same.

Any ideas why workers from Distributed (created from Slurm ClusterManagers) can’t use the GPUs on workers that clearly have them and is installed correctly?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [November 22, 2022, 7:32pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663/2 "2022-11-22T19:32:58Z")

</div>

> [@asaxton](#):
>
> ```julia
> Error showing value of type CUDA.DeviceIterator:
> ERROR: CUDA error (code 100, CUDA_ERROR_NO_DEVICE)
> 
> ```

this is probably trying to show your local CUDA device and you don’t have one. Can you try a simpler task that’s computational, something like

```julia
@fetechfrom procs[1] sum(CUDA.rand(10))

```

so that you’re using GPU but only transmitting “CPU-compatible” data over network

---

<div class="post-metadata">

**Author:** ![asaxton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/asaxton/32/22216_2.png) [@asaxton](https://discourse.julialang.org/u/asaxton)\
**Post date:** [November 22, 2022, 7:39pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663/3 "2022-11-22T19:39:13Z")

</div>

> [@jling](#):
>
> `@fetechfrom procs[1] sum(CUDA.rand(10))`

That worked! It returned `4.443556f0`. Sorry I didn’t think to try that in the first place.

So does this mean there’s a bug in `CUDA.devices()`? Any suggestions on how to document this better and report it?

---

<div class="post-metadata">

**Author:** ![asaxton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/asaxton/32/22216_2.png) [@asaxton](https://discourse.julialang.org/u/asaxton)\
**Post date:** [November 22, 2022, 7:39pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663/4 "2022-11-22T19:39:50Z")

</div>

P.S. I’m asking the system admin it install Julia 1.8.3 for me right now. Maybe the bug is already fixed?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [November 22, 2022, 8:27pm UTC](https://discourse.julialang.org/t/distributed-julia-clustermangers-package-slurm-and-cuda-on-a-shared-file-system/90663/5 "2022-11-22T20:27:36Z")

</div>

> [@asaxton](#):
>
> So does this mean there’s a bug in `CUDA.devices()`?

idk exactly but I imagine the problem is `CUDA.devices()` return pointers / reference to actual devices on the system it was called, when you transmit them back to your log in node, they’re wrong since the devices are not present on your login nodes, thus the error
