# Understanding memory usage in Julia

**URL:** <https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647>\
**Category:** Performance\
**Tags:** memory, memory-allocation\
**Created:** [July 15, 2023, 8:20am UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647 "2023-07-15T08:20:00Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 8:20am UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/1 "2023-07-15T08:20:00Z")

</div>

I am working on a finite element program that solves a non linear problem by incrementing the boundary conditions, that means that from time to time I need to assemble some quite large stiffness matrices (\approx 350k \times 350k) and solve linear algebra problems.

I am surprised of how much memory the program actually uses compared to what I expected and I am trying to understand why that is happening.

Before starting the main loop of the program, after the mesh has been read from the input file and all necessary data structure have been created I check the memory occupancy of all my variables with the @show\_local macro by @cgeoga from [this post](https://discourse.julialang.org/t/is-there-a-package-to-list-memory-consumption-of-selected-data-objects/85019/12) and I get a total usage of about 1.5gb, which is what I expected for a 113k elements model, but if I check the output from `pmap` on that process it says 30gb!

But that’s not all, as the calculations proceed the memory usage from `pmap` grows seemingly out of control reaching in excess of 100gb after two hundred increments, when it gets killed by OS

During the main loop the program needs to update a few stiffness matrices and solve linear problems, but everything happens inside functions, so temporary variables should be taken care of by the garbage collection, as no additional matrix or vectors are created but the existing updated.

I tried to force garbage collection invoking `GC.gc()` at the end of every load increment, but things didn’t change.

I am using `SparseArrays` to store the stiffness matrix and [LienearAlgebra.cholesky](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.cholesky) factorization to solve the linear algebra problems, I also store the cholesky factorization for re-use.

Everything runs on a node of a cluster, and I asked for two cores and 96gb,

can anyone help to understand what is happening and is it possible to contain the memory usage of a program of this type?

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [July 15, 2023, 11:49am UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/2 "2023-07-15T11:49:41Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> I am using `SparseArrays` to store the stiffness matrix and [LienearAlgebra.cholesky](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.cholesky) factorization to solve the linear algebra problems, I also store the cholesky factorization for re-use.

Are you using that function or the in-place version? [This one](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.cholesky!)

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 3:21pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/3 "2023-07-15T15:21:29Z")

</div>

I think what I am doing is equivalent to the one you pointed out, this is what I am doing:

```julia
cholB1KtB1ff = let 
  Ktu = makeKt(Φ, elems, u, d)
  B1KtB1ff = Symmetric(transpose(B1[:,ifree])*(λT/Δt^2*MM+λV*Ktu)*B1[:,ifree])
  cholesky(B1KtB1ff)
end

```

this should mean that neither `Ktu` nor `B1KtB1ff` are kept, and memory should be freed after the `end`, is that correct ?

would this be any different?

```julia
cholB1KtB1ff = let 
  Ktu = makeKt(Φ, elems, u, d)
  cholesky!(Symmetric(transpose(B1[:,ifree])*(λT/Δt^2*MM+λV*Ktu)*B1[:,ifree]))
end

```

where `MM` has the same dimensions of `Ktu`, is assembled once and never changes, and the other are scalars.

---

<div class="post-metadata">

**Author:** ![nsajko](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nsajko/32/221187_2.png) [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Post date:** [July 15, 2023, 3:54pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/4 "2023-07-15T15:54:03Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> I am surprised of how much memory the program actually uses compared to what I expected and I am trying to understand why that is happening.

If you want specific help it would probably be best if you included some self-contained code. As far as general tips go, check out the [Performance tips](https://docs.julialang.org/en/v1/manual/performance-tips/) section in the manual. In particular, try doing allocation profiling (very convenient in the (VS)Code editor, just do `@profview_allocs some_call()`).

The latest Julia release also has a feature, seemingly still not documented, to trigger frequent garbage collection after a certain threshold. For example: `julia --heap-size-hint=10G`.

> [@Andrea\_Vigliotti](#):
>
> I tried to force garbage collection invoking `GC.gc()`

If I remember correctly, sometimes it was necessary to invoke this call multiple times in a row.

> [@Andrea\_Vigliotti](#):
>
> this should mean that neither `Ktu` nor `B1KtB1ff` are kept, and memory should be freed after the `end`, is that correct ?

No, Julia isn’t C++. The memory gets freed when the garbage collector decides so.

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 4:09pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/5 "2023-07-15T16:09:03Z")

</div>

I think I am following all of the guidelines of [Performance tips](https://docs.julialang.org/en/v1/manual/performance-tips/), especially when it comes to enclosing code in functions, reusing vectors, and not having globals, I will look at `@profview_allocs some_call())` to see if I can get a better understanding of what is happening

About invoking `GC.gc()` multiple times, at the moment it is done at every increment of the boundary conditions, which is more often than the stiffness matrix factorization gets updated. But still the memory footprint of the program is from 50 to 100 times the expected

---

<div class="post-metadata">

**Author:** ![nsajko](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nsajko/32/221187_2.png) [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Post date:** [July 15, 2023, 4:11pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/6 "2023-07-15T16:11:49Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> I will look at `@profview_allocs some_call())` to see if I can get a better understanding of what is happening

Just remember to run with `@profview_allocs` an additional time before doing it _for real_, to get everything to compile.

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [July 15, 2023, 4:22pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/7 "2023-07-15T16:22:44Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> I also store the cholesky factorization for re-use.

The factorization objects (FO) returned by `cholesky` contain large arrays managed by the solver library (so not directly visible to Julia tooling). When an FO goes out of scope and is garbage-collected, the GC should call a finalizer which frees the underlying storage. Are you sure that your obsolete FO’s are going out of scope?

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 4:35pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/8 "2023-07-15T16:35:45Z")

</div>

I was thinking something like that was happening, but I was not aware of this peculiarity with the garbage collection of the FO objects,

in fact the memory occupancy of `cholB1KtB1ff ` from `Base.summarysize` is 8 bytes, while the size in memory of the matrix is \approx 150 mb, that suggested that the decomposition was handled and kept somewhere out of the reach of julia, but thought it was still handled by `GC.gc()` …

> [@Ralph\_Smith](#):
>
> the GC should call a finalizer which frees the underlying storage

can you suggest how to do it, or where could I find information about it?

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 4:43pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/9 "2023-07-15T16:43:28Z")

</div>

> [@Ralph\_Smith](#):
>
> The factorization objects (FO) returned by `cholesky` contain large arrays managed by the solver library

do you think it would make any difference if I switched to `MKL` instead of `OpenBLAS` that comes with the standard Julia installation?

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [July 15, 2023, 4:56pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/10 "2023-07-15T16:56:03Z")

</div>

Sorry if I wasn’t clear; the `gc` call should automatically free the memory by calling the finalizer without extra effort from you, as long as you don’t keep the `cholB1KtB1ff` variable (which is just a wrapped pointer) in scope. You might have been misled about the need to drop references to the FO’s by the small size visible to the Julia tools.

Finalizers are described in the Julia manual, and they are set up for the sparse array factorizations in the `SharedArrays` interface to the `SuiteSparse` library. Which BLAS library you use should not matter here.

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 5:11pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/11 "2023-07-15T17:11:24Z")

</div>

> [@Ralph\_Smith](#):
>
> Sorry if I wasn’t clear; the `gc` call should automatically free the memory by calling the finalizer without extra effort from you, as long as you don’t keep the `cholB1KtB1ff` variable (which is just a wrapped pointer) in scope.

I am keeping `cholB1KtB1ff` in the main function because I need it multiple times, and I overwrite it when I update it, what should I do to make sure that `gc()` takes care of disposing the old one when a new one is made?

> [@Ralph\_Smith](#):
>
> You might have been misled about the need to drop references to the FO’s by the small size visible to the Julia tools

I thought `cholB1KtB1ff` was a pointer to something living elsewhere, but I thought `gc` could reach it even if it was not visible through `Base.summarysize`

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [July 15, 2023, 5:12pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/12 "2023-07-15T17:12:32Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> `cholesky!(Symmetric(transpose(B1[:,ifree])*(λT/Δt^2*MM+λV*Ktu)*B1[:,ifree]))`

I meant reusing the input buffer, with:

```julia
cholesky!(A::AbstractMatrix, NoPivot(); check = true) -> Cholesky

```

```julia
The same as [`cholesky`](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.cholesky), but saves space by overwriting the input `A`, instead of creating a copy. An [`InexactError`](https://docs.julialang.org/en/v1/base/base/#Core.InexactError) exception is thrown if the factorization produces a number not representable by the element type of `A`, e.g. for integer types.

```

Also note that `A + B` will create a new Matrix as a result. If you are doing this iteratively, you probably should compute `C .= A .+ B` and reuse `C`. Matrix multiplications can be done in place with `mul!`.

And `B[:,ifree]` allocates a new vector. Use, perhaps, `@view(B1[:,ifree])`.

I don’t know how much each of these things might make a difference there for you.

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 5:23pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/13 "2023-07-15T17:23:44Z")

</div>

> [@lmiq](#):
>
> Also note that `A + B` will create a new Matrix as a result. If you are doing this iteratively, you probably should compute `C .= A .+ B` and reuse `C`

I was aware of that, but the old copies should be discharged by the `gc()`, right? I was not too much worried about that because it takes place only a limited number of times, and it is not where the bulk of the computations lie

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [July 15, 2023, 5:43pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/14 "2023-07-15T17:43:42Z")

</div>

Although what you describe seems reasonable to me, there may be some subtlety about scope and object references in play here. Instead of doing so much in your main function, perhaps you could have an intermediate function which does the factorization of a given matrix, applies it to all the vectors associated with that problem, and then returns, thus isolating each `cholB1KtB1ff`. Such function barriers sometimes help the memory management.

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 15, 2023, 5:46pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/15 "2023-07-15T17:46:26Z")

</div>

thanks, I’ll try that and I’ll post the results

---

<div class="post-metadata">

**Author:** ![danielwe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielwe/32/35657_2.png) [@danielwe](https://discourse.julialang.org/u/danielwe)\
**Post date:** [July 15, 2023, 7:43pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/16 "2023-07-15T19:43:40Z")

</div>

> [@lmiq](#):
>
> I meant reusing the input buffer, with:
> 
> ```julia
> cholesky!(A::AbstractMatrix, NoPivot(); check = true) -> Cholesky
> 
> ```

Note that OP is using sparse matrices, for which different conventions apply. You can do this:

```julia
F = cholesky(A)
# use F
cholesky!(F, B)
# use F

```

From the original post, it sounds like this is what OP is doing (“I also store the cholesky factorization for re-use”), but this is not what’s happening in the example code provided later.

I’m not an expert here, but in general, I think such reuse is only advisable if `A` and `B` have the same sparsity pattern. However, if that’s the case, it’s highly recommended, as it not only saves allocations but also saves you from recomputing the symbolic factorization.

> [@Andrea\_Vigliotti](#):
>
> do you think it would make any difference if I switched to `MKL` instead of `OpenBLAS`

BLAS libraries don’t matter here. Sparse linear algebra is handled by different libraries altogether. Cholesky factorization and solving linear systems using the factorization object are both done by SuiteSparse.CHOLMOD, and matrix-vector multiplication is handled by a Julia native function within the SparseArray stdlib.

* * *

I can’t really help with the memory issue. I’ve also observed that memory will sometimes grow indefinitely in long-running code with lots of sparse arrays, so I’ve been suspecting that the wrapped memory underlying `SparseMatrixCSC` isn’t always properly finalized, but I haven’t made any attempt to find the bug. This issue might be relevant: [Memory leak in solving ODE many times · Issue #1946 · SciML/OrdinaryDiffEq.jl · GitHub](https://github.com/SciML/OrdinaryDiffEq.jl/issues/1946).

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [July 15, 2023, 8:51pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/17 "2023-07-15T20:51:44Z")

</div>

> [@Andrea\_Vigliotti](#):
>
> Everything runs on a node of a cluster, and I asked for two cores and 96gb

Is your Julia program parallel, i.e. with two tasks (via @spawn or @threads)? Julia’s gc is not very good at managing memory in such cases. Actually, the only times I’ve had problems with memory usage growing without bounds have been with allocation in multiple threads.

---

<div class="post-metadata">

**Author:** ![Andrea\_Vigliotti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_vigliotti/32/7635_2.png) [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Post date:** [July 16, 2023, 7:12am UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/19 "2023-07-16T07:12:09Z")

</div>

> [@sgaure](#):
>
> Is your Julia program parallel, i.e. with two tasks (via @spawn or @threads)?

yes, the code is parallel and uses `@threads` in a `for` loop for assembling the global stiffness matrix from the elemental stiffness matrices, but no memory is allocated within the `for` loop, all the required arrays are allocated in advance. I will try to remove that and report what happens.

> [@danielwe](#):
>
> From the original post, it sounds like this is what OP is doing (“I also store the cholesky factorization for re-use”), but this is not what’s happening in the example code provided later.

what I meant by re-using is that the same factorization is used for a few different problems before a new stiffness matrix is assembled and factorized. That is because the cost for factorization is the highest and it makes sense to keep using the same factorized matrix at the cost of taking more iterations rather than assembling and factorizing at each step (that would take one iteration only)

reusing the factorization in the sense you showed in your post proved not to be effective in my case

---

<div class="post-metadata">

**Author:** ![pjuergens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pjuergens/32/28380_2.png) [@pjuergens](https://discourse.julialang.org/u/pjuergens)\
**Post date:** [July 24, 2023, 9:47am UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/20 "2023-07-24T09:47:42Z")

</div>

I’m facing the same problem, i.e. memory usage growing without bounds and I need to use multiple cores with allocations. How did you solve the problem in that case? @Andrea_Vigliotti did you test what happened when you removed parallelization?

In my case, parallelization is done via `@everywhere` and `pmap`. (and no `SparseArrays`)

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [July 24, 2023, 4:21pm UTC](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647/21 "2023-07-24T16:21:05Z")

</div>

I added:

```julia
if Sys.free_memory()/2^30 < 6.0
    GC.gc()
end

```

to large functions that are called by pmap…

I think there is a bug in the garbage collector, this should not be needed…

Created a bug report: [Garbage collector not working · Issue #50658 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/50658)

[Next page](https://discourse.julialang.org/t/understanding-memory-usage-in-julia/101647.md?page=2)
