# Initializing local variables on workers without using excessive memory

**URL:** <https://discourse.julialang.org/t/initializing-local-variables-on-workers-without-using-excessive-memory/114713>\
**Category:** Julia at Scale\
**Tags:** jump, distributed, optimization\
**Created:** [May 25, 2024, 11:08am UTC](https://discourse.julialang.org/t/initializing-local-variables-on-workers-without-using-excessive-memory/114713 "2024-05-25T11:08:46Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![lgo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lgo/32/48751_2.png) [@lgo](https://discourse.julialang.org/u/lgo)\
**Post date:** [May 25, 2024, 11:08am UTC](https://discourse.julialang.org/t/initializing-local-variables-on-workers-without-using-excessive-memory/114713/1 "2024-05-25T11:08:46Z")

</div>

I’m using Julia and JuMP to solve a massive optimization problem with decomposition on a HPC. For this purpose, I want to initialize an independent optimization problem on each worker and then solve it repeatedly. So far, my setup to initialize the local problem on each worker corresponds to the minimum example below. In the actual problem, `b` is large and requires substantial resources to create, so it’s a step that needs to be done locally.

```julia
using Distributed
addprocs(2)

# initialize function on worker
@everywhere begin 
    testFunc(b) = b * 2
end

# initialize local variable on each worker
w = []
for x in workers()
    ini = @async @everywhere x begin
        b = myid()
    end
    push!(w, ini)
end
wait.(w)

# call function on worker
fetch(Distributed.@spawnat 3 testFunc(b))

```

This setup works on small-scale examples, but the initialization step requires excessive memory. When I benchmarked the initialization step for a large-scale example separately without distributed computing, it never required more than 22 GB. However, in the distributed case, my workers always run out of memory and get killed, although they have 32 GB available, and I use the `heap-size-hint` flag with a value of 28 GB.

I noticed that the initialization does not run out of memory if I replace `@async @everywhere` in the example above with `@spawnat`. However, the variable seems to be initialized in the wrong scope then, and the subsequent function call fails. Yet, it shows that the problem is not with the memory allocation for the workers.

---

<div class="post-metadata">

**Author:** ![lgo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lgo/32/48751_2.png) [@lgo](https://discourse.julialang.org/u/lgo)\
**Post date:** [May 29, 2024, 10:55am UTC](https://discourse.julialang.org/t/initializing-local-variables-on-workers-without-using-excessive-memory/114713/2 "2024-05-29T10:55:48Z")

</div>

I figured, I can use `@spawnat` but have to make `b` a global variable, so it is accessible outside of the function scope. I’m still curios why @everywhere combined with a specific worker allocation uses so much memory.

```julia
using Distributed
addprocs(2)

# initialize function on worker
@everywhere begin 
    testFunc(b) = b * 2
end

# initialize local variable on each worker
w = []
for x in workers()
    ini = @spawnat x begin
        global b = myid()
    end
    push!(w, ini)
end
wait.(w)

# call function on worker
fetch(Distributed.@spawnat 3 testFunc(b))

```
