# Precompilation gridlock on HPC cluster

**URL:** <https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213>\
**Category:** General Usage\
**Tags:** question\
**Created:** [May 14, 2024, 2:28am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213 "2024-05-14T02:28:57Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [May 14, 2024, 2:28am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/1 "2024-05-14T02:28:57Z")

</div>

I’m running a batch script on an HPC cluster where each Julia execution is expected to be \<1 min, within a bash loop. I ran into this weird “precompilation gridlock”, shown below. Has anybody experienced something like this?

```julia
Precompiling MLJBase
  Progress [=======================================>] 35/36
  ✓ FixedPointNumbers
  ✓ ColorTypes
  ? Distributions → DistributionsChainRulesCoreExt
  ◐ MLJBase Being precompiled by another machine (hostname: worker6062, pid: 603518, pidfile: /mnt/home/mcranmer/.julia/compiled/v1.10/MLJBase/jaWQl…

```

Basically it looks like all the workers wait for one worker to finish precompiling. Then when that worker finally finishes\[1\], the next one decides that the precompilation cache was invalidated, and it needs to precompile again. This process repeats over and over.

The result is that out of 3200 cores across the cluster, only 1 is ever in use, since precompilation ends up taking longer the processing itself:

 ![Screenshot 2024-05-14 at 03.23.34](https://global.discourse-cdn.com/julialang/original/3X/5/f/5fa98ddecbc6103eda7fdb6d8b1b0e41532cb6af.png)

How can I prevent this? Or is there a way I can force each worker to avoid waiting on another process to finish precompilation?

Alternatively, is there a way to disable precompilation altogether, so that only “compilation” occurs?

* * *

1. Which takes forever as this machine has many very slow cores.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [May 14, 2024, 2:51am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/2 "2024-05-14T02:51:28Z")

</div>

I might be jumping to conclusion a bit but at first glance it looks like a variant of:

- [Error precompiling on cluster](https://discourse.julialang.org/t/error-precompiling-on-cluster/104974/)
- [Precompilation Fails on HPC](https://discourse.julialang.org/t/precompilation-fails-on-hpc/104664)
- [Yet another precompilation-on-HPC issue](https://discourse.julialang.org/t/yet-another-precompilation-on-hpc-issue/105731/)

I think they started to pop up in 1.9 because that’s when we tightened the criteria for invalidation (i.e. more easily invalid)

and the solution is to play with `JULIA_CPU_TARGET` [`compilecache` failed when `@everywhere using` from remote machines · Issue #48217 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/48217#issuecomment-1554149169)

---

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [May 14, 2024, 2:55am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/3 "2024-05-14T02:55:34Z")

</div>

Thanks.

One other clue is that this gridlock only started when I tried to interact with the environment from another interactive node to visualize some stuff. That seemed to invalidate the cache (maybe due to `-O2` vs `-O3`). After that point the gridlock started (it’s a shared filesystem, so the cache would be shared by both my workers and interactive REPL).

That seemed to make the workers go into a loop where they would keep invalidating the cache of the previous one.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [May 14, 2024, 2:59am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/4 "2024-05-14T02:59:28Z")

</div>

yeah, I want to add that IMO it’s overly tight – in the github issue (last link above), the login and remote nodes are using the same CPUs but one has 2 NUMA nodes the other has only 1, but I don’t believe that should have changed the compile cache hash

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [May 14, 2024, 3:45am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/5 "2024-05-14T03:45:33Z")

</div>

You have a heterogeneous cluster and your processes are overwriting the precompation of the other.

You might want to use a distinct `JULIA_DEPOT` for your visualization node or somehow set `JULIA_CPU_TARGET` appropriately.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [May 14, 2024, 3:56am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/6 "2024-05-14T03:56:19Z")

</div>

I don’t know what’s going on because I never got to work with clusters, but I could search relevant sections in the docs from some keywords here.

- [Frequently Asked Questions · The Julia Language](https://docs.julialang.org/en/v1/manual/faq/#Computing-cluster)
- [Environment Variables · The Julia Language](https://docs.julialang.org/en/v1/manual/environment-variables/#JULIA_CPU_TARGET)
- [System Image Building · The Julia Language](https://docs.julialang.org/en/v1/devdocs/sysimg/#sysimg-multi-versioning)

---

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [May 14, 2024, 11:08am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/7 "2024-05-14T11:08:52Z")

</div>

Just to note; my cluster is **not** heterogenous. The node I am visualising things from is the same type as the nodes I am running from.

My guess is that it’s either:

- My login shell prescribes different Julia env variables (like `-O2`) which triggered a re-precompilation, (and thus caused workers to get stuck while waiting for it to finish) or
- I added a package while the batch script was already running.

The weird thing is even after one worker pre-compiled, the other ones started precompiling again, one after the other, (even though they are identical worker nodes). I don’t understand that. Maybe it’s about where they were in the precompilation process the moment the global mutable cache got changed, and so their new precompilation layer was somehow immediately invalid?

In either case I’d like to figure out how to prevent this. Is there a way I can freeze the precompilation cache when I execute my job, so that it doesn’t interact with a global mutable cache? Or, just turn off precompilation?

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [May 14, 2024, 12:22pm UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/8 "2024-05-14T12:22:19Z")

</div>

@ianshmean is the pidlock per package or per file? The later includes in the hash the optimization flags, but the former would maybe cause this issue.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [May 14, 2024, 1:19pm UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/9 "2024-05-14T13:19:01Z")

</div>

`ENV["JULIA_DEBUG"] ="loading"` may help debug why it is invalidating the cache.

---

<div class="post-metadata">

**Author:** ![ianshmean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ianshmean/32/216042_2.png) [@ianshmean](https://discourse.julialang.org/u/ianshmean)\
**Post date:** [May 16, 2024, 4:12am UTC](https://discourse.julialang.org/t/precompilation-gridlock-on-hpc-cluster/114213/10 "2024-05-16T04:12:26Z")

</div>

> [@vchuravy](#):
>
> [@ianshmean](https://discourse.julialang.org/u/ianshmean) is the pidlock per package or per file?

It’s specific to everything that the resulting hash in the cache filename is, except:

- ignores the active project, so two processes with different projects will only result in one doing work
- ignores preferences because they cannot be hashed before spawning the precompilation process

> <https://github.com/JuliaLang/julia/blob/4980544ce1f3f85f22ac801d47c7c94bf5764f15/base/loading.jl#L3586-L3589>
