# How to debug remote worker not connecting?

**URL:** <https://discourse.julialang.org/t/how-to-debug-remote-worker-not-connecting/20762>\
**Category:** General Usage\
**Created:** [February 13, 2019, 7:54pm UTC](https://discourse.julialang.org/t/how-to-debug-remote-worker-not-connecting/20762 "2019-02-13T19:54:13Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [February 13, 2019, 7:54pm UTC](https://discourse.julialang.org/t/how-to-debug-remote-worker-not-connecting/20762/1 "2019-02-13T19:54:13Z")

</div>

I’m using ClusterManagers.jl to create an ElasticManager to connect some workers on a cluster. This is working between some nodes on the cluster (manager and workers both on login nodes) but not between others (manager on login node, workers on compute nodes, which is in fact what I need). I’m guessing this is signaling some sort of networking/firewall problems.

Here’s what I’ve found:

- I can `telnet` between compute/login nodes on the same ports I use for the Julia cluster just fine.
- I can by-hand create a Julia `Socket` connection between compute/login nodes and send data back-and-forth just fine.
- Creating the ElasticManager and connecting a remote worker does _not_ work. The series of commands is I run `ElasticManager(...)` on the login node then from the compute nodes I do `julia -e 'using ClusterManagers; ClusterManagers.elastic_worker(<cookie>,<ip>,<port>)` to connect the workers, but this times out after 60sec. I’ve got it traced to that the connection process at least makes it to calling and returning from the `launch` function at [https://github.com/JuliaLang/julia/blob/v1.1.0/stdlib/Distributed/src/cluster.jl#L399](https://github.com/JuliaLang/julia/blob/v1.1.0/stdlib/Distributed/src/cluster.jl#L399) but I don’t have a good way to insert debug statements in the remaining Julia stdlib code itself (short of a slow-to-iterate Julia recompile).

Any suggestions how I can figure out what’s happening and fix it? Thanks.

---

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [February 13, 2019, 9:20pm UTC](https://discourse.julialang.org/t/how-to-debug-remote-worker-not-connecting/20762/2 "2019-02-13T21:20:27Z")

</div>

Ok, I think I’ve found the cause of the issue. Seems the cluster blocks TCP connections from the login nodes to the compute nodes. I thought Julia only needed the other direction to be allowed, since the worker processes on the compute nodes connect to the login node, but after digging into the code a bit it seems like after the initial connection, the master process then initiates a connection in the other direction ([https://github.com/JuliaLang/julia/blob/v1.1.0/stdlib/Distributed/src/managers.jl#L437](https://github.com/JuliaLang/julia/blob/v1.1.0/stdlib/Distributed/src/managers.jl#L437)), and its this that’s failing. Not sure a good workaround here…

(Btw, for anyone wondering, turns out good way to insert debug statements into Julia Stdlib or Base without recompiling Julia is just to `@eval` them directly into these modules)
