# Unreasonable memory usage with M4 GPU

**URL:** https://discourse.julialang.org/t/unreasonable-memory-usage-with-m4-gpu/123890
**Category:** GPU
**Tags:** metaljl
**Created:** [December 16, 2024, 2:42pm UTC](https://discourse.julialang.org/t/unreasonable-memory-usage-with-m4-gpu/123890 "2024-12-16T14:42:44Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![Andrea\_Pagnani](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_pagnani/32/3136_2.png) [@Andrea\_Pagnani](https://discourse.julialang.org/u/Andrea_Pagnani)
#### Post date: [December 16, 2024, 2:42pm UTC](https://discourse.julialang.org/t/unreasonable-memory-usage-with-m4-gpu/123890/1 "2024-12-16T14:42:44Z")

</div>

Dearests,

my fight to make a reasonable use of my M4 GPU continues.

> **Metal.versioninfo()**
>
> macOS 15.1.1, Darwin 24.1.0
> 
> Toolchain:
> 
> - Julia: 1.11.2
> - LLVM: 16.0.6
> 
> Julia packages:
> 
> - Metal.jl: 1.4.2
> - GPUArrays: 10.3.1
> - GPUCompiler: 0.27.8
> - KernelAbstractions: 0.9.31
> - ObjectiveC: 3.1.0
> - LLVM: 9.1.3
> - LLVMDowngrader\_jll: 0.3.0+2
> 
> 1 device:
> 
> - Apple M4 Pro (48.953 MiB allocated)

I developed a simple optimization problem (more of an MWE than what I need to do). I observe an explosion in memory. Before giving you the not-so-minimal example let me explain what I see.

```julia
function trainmodel!(model::Model; nepochs=100, verbose=true)
    opt = Flux.setup(Flux.Optimisers.Adam(0.1), model)
    for it in 1:nepochs
        grads = Flux.gradient(model) do m
            TestMetal.losslinearalgebra(m)
        end
        verbose && println("it = $it |grad| = $(norm(grads[1].msa))")
        Flux.Optimise.update!(opt, model, grads[1])
        GC.gc()
    end
end

```

The crux of the problem is that without the `GC.gc()` command after `update!`, the memory explodes when I use `MtlArray` Arrays (explode = computer becomes unresponsive for large memory usage). For normal Arrays, there is no problem.

If you want to run the full thing is a bit complicated, but doable. I created the gist below:

> <https://gist.github.com/pagnani/b9168ba36c0a2a5ac897faeb6b965056>

To use it you should

```julia
julia> include("testmetal.jl"); using .TestMetal
julia> q,L,M = 21,53,10_000; modelgpu=TestMetal.Model(q,L,M,gpu=true); modelcpu=TestMetal.Model(q,L,M,gpu=false);
julia> TestMetal.trainmodel!(modelgpu,nepochs=100,verbose=true) # beware that this is where my computer becomes unresponsive

```

Worth reporting upstream to Metal?  
Thanks  
A

---

<div class="post-metadata">

### Author: ![pxl-th](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pxl-th/32/31939_2.png) [@pxl-th](https://discourse.julialang.org/u/pxl-th)
#### Post date: [December 20, 2024, 5:30pm UTC](https://discourse.julialang.org/t/unreasonable-memory-usage-with-m4-gpu/123890/2 "2024-12-20T17:30:05Z")

</div>

You may keep an eye for caching allocator then 🙂

> <https://github.com/JuliaGPU/GPUArrays.jl/pull/576>
>
> Since Julia's GC is not aware of GPU memory, in scenarious with lots of allocati…ons we end up in either OOM situations or in excessively high memory usage.
> Even though the program may require only fraction of it.
> 
> To help with GPU memory utilizaton in a program with repeating blocks of code, we can wrap those regions in a scope that will utilize caching allocator every time the program enters this scope during execution.
> 
> For example, this is especially useful when training models, where you compute loss, gradients w.r.t. loss and perform in-place parameter update of the model.
> 
> \`\`\`julia
> model = ...
> for i in 1:1000
> GPUArrays.@cache\_scope kab :loop begin
> loss, grads = ...
> update!(optimizer, model, grads)
> end
> end
> \`\`\`
> 
> The caching allocator is defined by its name and is per-device (it will use current TLS device).
> 
> \### Example
> 
> In the following example we apply caching allocator at every iteration of the for-loop.
> Every iteration requires 2 GiB of gpu memory, without caching allocator
> GC wouldn't be able to free arrays in time resulting in higher memory usage.
> With caching allocator, memory usage stays at exactly 2 GiB.
> 
> After the loop, we free all cached memory if there's any (e.g. CUDA.jl will bulk-free immediately after execution of expression inside \`@cache\_scope\`, because it has performant allocator).
> 
> \`\`\`julia
> kab = CUDABackend()
> n = 1024^3
> CUDA.@sync for i in 1:1000
> GPUArrays.@cache\_scope kab :loop begin
> sin.(CUDA.rand(Float32, n))
> end
> end
> GPUArrays.invalidate\_cache\_allocator!(kab, :loop)
> \`\`\`
> 
> \### Backend differences
> 
> \- Because CUDA has more performant allocator, CUDA.jl will bulk-free arrays at the end of \`expr\` execution, instead of caching the arrays (\`free\_immediately=true\`).
> \- AMDGPU.jl instead caches them (\`free\_immediately=false\`) until user invalidates the cache.
> 
> \### Performance impact
> 
> Executing \[GaussianSplatting.jl\](https://github.com/JuliaNeuralGraphics/GaussianSplatting.jl/pull/26) benchmark (1k training iterations) on RX 7900XTX:
> 
> ||Without caching allocator|With caching allocator|
> |-|-|-|
> |GPU memory utilization|!\[image\](https://github.com/user-attachments/assets/92e0e802-9784-4e5a-85d0-7dfa7b4e8dbf)|!\[image\](https://github.com/user-attachments/assets/ab4c68eb-2caa-4a56-9c0a-9176285ef66d)|
> |Time|\`59.656476\` seconds|\`46.365646\` seconds|
> 
> \### TODO
> 
> \- \[x\] Support for 1.10.
> \- \[x\] Support bulk-freeing instead of caching.
> \- \[x\] Add PR description.
> \- \[x\] Documentation.
> \- \[x\] Tests.
> 
> \### PRs for other GPU backends
> 
> \- AMDGPU: https://github.com/JuliaGPU/AMDGPU.jl/pull/710
> \- CUDA: https://github.com/JuliaGPU/CUDA.jl/pull/2593

---

<div class="post-metadata">

### Author: ![Andrea\_Pagnani](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrea_pagnani/32/3136_2.png) [@Andrea\_Pagnani](https://discourse.julialang.org/u/Andrea_Pagnani)
#### Post date: [December 21, 2024, 8:48am UTC](https://discourse.julialang.org/t/unreasonable-memory-usage-with-m4-gpu/123890/3 "2024-12-21T08:48:08Z")

</div>

Thx! I’ll keep an eye to your PR
