# HPC cluster SSH security issue? (password vs passwordless ssh keypair)

**URL:** <https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664>\
**Category:** General Usage\
**Tags:** hpc\
**Created:** [January 20, 2021, 11:40am UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664 "2021-01-20T11:40:04Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![henrikjaerleblad](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrikjaerleblad/32/21172_2.png) [@henrikjaerleblad](https://discourse.julialang.org/u/henrikjaerleblad)\
**Post date:** [January 20, 2021, 11:40am UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/1 "2021-01-20T11:40:04Z")

</div>

Hello,

Previously, I used to be able to run multi-core jobs on my university’s HPC SLURM cluster like this:

_#!/bin/bash_  
_#SBATCH --mail-type=ALL_  
_#SBATCH --mail-user=mymail@mailymail.country_  
_#SBATCH --export=ALL_  
_#SBATCH -J a\_julia\_job_  
_#SBATCH --partition=xeon16_  
_#SBATCH -n 64_  
_#SBATCH -N 4_  
_#SBATCH --mem=50G_  
_#SBATCH --time=5-00:00:00_  
_#SBATCH --output=log\_file\_good\_info.out_  
_srun hostname | sort \> nodefile.$SLURM\_JOBID_  
_julia --machine-file nodefile.$SLURM\_JOBID /the/path/to/the/script/thascript.jl inputs\_file.jl_

But now, I get an authentication error message after a couple of seconds, even when I try to use only my main computational node (instead of 4 in total, for example):

_Permission denied, please try again.^M_  
_Permission denied, please try again.^M_  
_Received disconnect from 12.3.456.7 port 22:2: Too many authentication failures^M_  
_Authentication failed.^M_  
_ERROR: TaskFailedException:_  
_Unable to read host:port string from worker. Launch command exited with error?_  
_Stacktrace:_  
_[1] worker\_from\_id(::Distributed.ProcessGroup, ::Int64) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:1059_  
_[2] worker\_from\_id at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:1056 [inlined]_  
_[3] #remote\_do#156 at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/remotecall.jl:482 [inlined]_  
_[4] remote\_do at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/remotecall.jl:482 [inlined]_  
_[5] kill at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/managers.jl:534 [inlined]_  
_[6] create\_worker(::Distributed.SSHManager, ::WorkerConfig) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:581_  
_[7] setup\_launched\_worker(::Distributed.SSHManager, ::WorkerConfig, ::Array{Int64,1}) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:523_  
_[8] (::Distributed.var"#43#46"{Distributed.SSHManager,Array{Int64,1},WorkerConfig})() at ./task.jl:333_  
_Stacktrace:_  
_[1] sync\_end(::Array{Any,1}) at ./task.jl:300_  
_[2] macro expansion at ./task.jl:319 [inlined]_  
_[3] #addprocs\_locked#40(::Base.Iterators.Pairs{Symbol,Any,Tuple{Symbol,Symbol,Symbol},NamedTuple{(:tunnel, :sshflags, :max\_parallel),Tuple{Bool,Cmd,Int64}}}, ::typeof(Distributed.addprocs\_locked), ::Distributed.SSHManager) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:477_  
_[4] #addprocs\_locked at ./none:0 [inlined]_  
_[5] #addprocs#39(::Base.Iterators.Pairs{Symbol,Any,Tuple{Symbol,Symbol,Symbol},NamedTuple{(:tunnel, :sshflags, :max\_parallel),Tuple{Bool,Cmd,Int64}}}, ::typeof(addprocs), ::Distributed.SSHManager) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:441_  
_[6] #addprocs at ./none:0 [inlined]_  
_[7] #addprocs#243(::Bool, ::Cmd, ::Int64, ::Base.Iterators.Pairs{Union{},Union{},Tuple{},NamedTuple{(),Tuple{}}}, ::typeof(addprocs), ::Array{Any,1}) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/managers.jl:118_  
_[8] addprocs(::Array{Any,1}) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/managers.jl:117_  
_[9] process\_opts(::Base.JLOptions) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.3/Distributed/src/cluster.jl:1305_  
_[10] #invokelatest#1 at ./essentials.jl:709 [inlined]_  
_[11] invokelatest at ./essentials.jl:708 [inlined]_  
_[12] exec\_options(::Base.JLOptions) at ./client.jl:254_  
_[13] \_start() at ./client.jl:460_

I believe I ran a Pkg.update() shortly before this. Nevertheless, do you think this is an error caused by the Julia update, a cluster security update of some kind or a fault of my own? I have double-checked so that I haven’t made any errors, but of course I can have overlooked something still.

I am aware that the Julialang FAQs recommend running Julia scripts with options with #!/bin/bash in a different way. Maybe that is what I should do? However, this way of doing it has worked perfectly fine on my HPC cluster up until this week.

Grateful for any thoughts/comments! Sorry if this post/question should go somewhere else/stack overflow. If that is the case, please re-direct me.

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [January 20, 2021, 12:14pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/2 "2021-01-20T12:14:19Z")

</div>

Looks like a security update to me.  
Slurm clusters use something called munge for authentication. It could be that the cluster admins have disabled ssh access.

TO check I would write a quick jobscript which produces that nodefile  
Then write a bach loop which runs trhoug the list and does  
ssh $host date

Or just run a two job interactive Slurm job and try to ssh between the compute nodes

I should say I am basing my response on those first 4 lines, but it is something you should be able to check quickly

---

<div class="post-metadata">

**Author:** ![henrikjaerleblad](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrikjaerleblad/32/21172_2.png) [@henrikjaerleblad](https://discourse.julialang.org/u/henrikjaerleblad)\
**Post date:** [January 20, 2021, 12:18pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/3 "2021-01-20T12:18:57Z")

</div>

Thank you. Unfortunately on my university cluster, interactive Slurm jobs are disabled. But I am starting to think this is a cluster security issue for sure.

Update: I should now add that I am investigating whether it might have something to do with me recently creating an ssh keypair with password, rather than using the standard non-password ssh keypair. Does anyone know if Julia looks at an ssh key, and examines whether there is a required password or not? And then tries to connect accordingly?

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [January 20, 2021, 12:21pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/4 "2021-01-20T12:21:29Z")

</div>

You should be able to use an ssh agent in the job though - actually I don;t have experience with that as I would always use a passwordless pair on a cluster.

I guess the other solution would be a passwordless pair and some sort of .ssh/config which says ‘only use this passwordless air when using cluster nodes’  
I think this is quite easy

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [January 20, 2021, 12:27pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/5 "2021-01-20T12:27:46Z")

</div>

I am going to be shot for this advice - by security mavens  
The .ssh/config could read

```julia
Host node*
    IdentityFile ~/.ssh/my-key-with-no-password

I think configuring ssh-agent in the job may be better - and keep the passphrase for the key in a file in your home directory, don't expose it in the job script

```

---

<div class="post-metadata">

**Author:** ![henrikjaerleblad](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrikjaerleblad/32/21172_2.png) [@henrikjaerleblad](https://discourse.julialang.org/u/henrikjaerleblad)\
**Post date:** [January 20, 2021, 12:47pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/6 "2021-01-20T12:47:48Z")

</div>

Yes that could be a way to go. I agree, not exposing it in the job script would be preferable, and make the world a safer place.

Thanks. I however also have the possibility of going back to only using a passwordless keypair. I will think a little bit about if I really need a password keypair at the moment. It was something I created just because anyway.

Another update: It works again. The .ssh/authorized\_keys file was re-generated and that solved the issue.

---

<div class="post-metadata">

**Author:** ![Alexander-Barth](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alexander-barth/32/3692_2.png) [@Alexander-Barth](https://discourse.julialang.org/u/Alexander-Barth)\
**Post date:** [January 20, 2021, 1:21pm UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/7 "2021-01-20T13:21:09Z")

</div>

> [@henrikjaerleblad](#):
>
> julia --machine-file nodefile …

I had a similar problem in the past. If I remember correctly, if you call julia like this you bypass SLURM (and the accounting of computing resources) and depending on the cluster configuration this might not be allowed or recommended. I ended up using `ClusterManagers` which integrates nicely with SLURM and uses `srun` to launch the processes on the worker node ([https://github.com/JuliaParallel/ClusterManagers.jl/blob/master/src/slurm.jl#L54](https://github.com/JuliaParallel/ClusterManagers.jl/blob/master/src/slurm.jl#L54)) .

In my SLURM script, I have code block like this (where `$script` is the full path of the julia script):

```julia
julia <<EOF
using Distributed
using ClusterManagers
addprocs(SlurmManager($SLURM_NTASKS))

hosts = []
pids = []
for i in workers()
    host, pid = fetch(@spawnat i (gethostname(), getpid()))
    push!(hosts, host)
    push!(pids, pid)
end

@show hosts

include("$script")

for i in workers()
    rmprocs(i)
end
EOF

```

You can also skip the `gethostname(), getpid()` part, but it can be useful for troubleshooting.  
No SSH logins are necessary for `ClusterManagers.SlurmManager`.

---

<div class="post-metadata">

**Author:** ![henrikjaerleblad](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrikjaerleblad/32/21172_2.png) [@henrikjaerleblad](https://discourse.julialang.org/u/henrikjaerleblad)\
**Post date:** [January 22, 2021, 7:21am UTC](https://discourse.julialang.org/t/hpc-cluster-ssh-security-issue-password-vs-passwordless-ssh-keypair/53664/8 "2021-01-22T07:21:16Z")

</div>

That was a really good example of a better way to do it. I tried it just now and it worked perfectly. In addition to being the correct way of doing it and giving more log for de-bugging, I also think it creates much less memory overhead. Need to investigate this further though.

Thank you very much!
