# Unexpected OOM errrors in julia 1.9.0 and 1.9.1 with Distributed

**URL:** <https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103>\
**Category:** Julia at Scale\
**Tags:** memory-allocation\
**Created:** [June 9, 2023, 12:53pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103 "2023-06-09T12:53:53Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![vfonov](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vfonov/32/31828_2.png) [@vfonov](https://discourse.julialang.org/u/vfonov)\
**Post date:** [June 9, 2023, 12:53pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/1 "2023-06-09T12:53:54Z")

</div>

I have a relatively simple script that used to run fine on 1.8.X version: it executes lots of relatively small jobs via pmap.  
Executing on julia 1.8.5 with -p 32 , creates 32 processes , each using approximately 2Gb and after a few hours the calculation successfully finishes. However on 1.9.0 and 1.9.1 after a few minutes each julia processes starts to allocate lots of memory and once it reaches about 10Gb per task, all of them gets killed by oom-killer.  
Reducing number of tasks, by making each one of them process more data via internal loop does not seem to help.

Potentially this is related to [Garbage collection not aggressive enough on Slurm Cluster](https://discourse.julialang.org/t/garbage-collection-not-aggressive-enough-on-slurm-cluster/61649) , although my jobs seem to run fine on 1.8.X

---

<div class="post-metadata">

**Author:** ![sloede](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sloede/32/44787_2.png) [@sloede](https://discourse.julialang.org/u/sloede)\
**Post date:** [June 9, 2023, 2:26pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/2 "2023-06-09T14:26:08Z")

</div>

We have encountered similar issues when during CI testing for [Trixi.jl](https://github.com/trixi-framework/Trixi.jl). What helped us in the parallel case was to add a `--heap-size-hint=1G` to each invocation of Julia, such that the garbage collector is more aggressive:

> <https://github.com/trixi-framework/Trixi.jl/blob/c47b6f6ae038535d04318c3294ff3f4a4cc41d11/test/runtests.jl#L30>

This actually helped to resolve some issues we had with parallel runs in a CI job before with v1.8, thus maybe it will help you as well.

---

<div class="post-metadata">

**Author:** ![vfonov](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vfonov/32/31828_2.png) [@vfonov](https://discourse.julialang.org/u/vfonov)\
**Post date:** [June 9, 2023, 2:38pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/3 "2023-06-09T14:38:24Z")

</div>

I tried running my script with --heap-size-hint=2G, but it didn’t change anything. Perhaps, this flag doesn’t propagate to julia processes that are started with `-p` flag?

---

<div class="post-metadata">

**Author:** ![sloede](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sloede/32/44787_2.png) [@sloede](https://discourse.julialang.org/u/sloede)\
**Post date:** [June 9, 2023, 2:45pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/4 "2023-06-09T14:45:55Z")

</div>

> [@vfonov](#):
>
> Perhaps, this flag doesn’t propagate to julia processes that are started with `-p` flag?

Good question. I have never used Distributed before, only MPI-based parallelism.

---

<div class="post-metadata">

**Author:** ![vfonov](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vfonov/32/31828_2.png) [@vfonov](https://discourse.julialang.org/u/vfonov)\
**Post date:** [June 9, 2023, 3:06pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/5 "2023-06-09T15:06:03Z")

</div>

Explicitly calling GC.gc() at the end of each parallel job, seem to have solved the problem - still running after 15min without OOM.

---

<div class="post-metadata">

**Author:** ![nathan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nathan/32/209855_2.png) [@nathan](https://discourse.julialang.org/u/nathan)\
**Post date:** [September 15, 2023, 8:36pm UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/6 "2023-09-15T20:36:30Z")

</div>

Is this the same as [OOM despite `--heap-size-hint` · Issue #50658 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/50658) ?

---

<div class="post-metadata">

**Author:** ![vfonov](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vfonov/32/31828_2.png) [@vfonov](https://discourse.julialang.org/u/vfonov)\
**Post date:** [September 28, 2023, 2:38am UTC](https://discourse.julialang.org/t/unexpected-oom-errrors-in-julia-1-9-0-and-1-9-1-with-distributed/100103/7 "2023-09-28T02:38:18Z")

</div>

Looks like it.
