# Significant CUDA.jl memory allocations outside of main pool?

**URL:** <https://discourse.julialang.org/t/significant-cuda-jl-memory-allocations-outside-of-main-pool/59991>\
**Category:** GPU\
**Tags:** memory\
**Created:** [April 25, 2021, 2:28pm UTC](https://discourse.julialang.org/t/significant-cuda-jl-memory-allocations-outside-of-main-pool/59991 "2021-04-25T14:28:24Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jpdoane](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jpdoane](https://discourse.julialang.org/u/jpdoane)\
**Post date:** [April 25, 2021, 2:28pm UTC](https://discourse.julialang.org/t/significant-cuda-jl-memory-allocations-outside-of-main-pool/59991/1 "2021-04-25T14:28:24Z")

</div>

I have been benchmarking various CUDA.jl kernels with large CuArrays, and have been experiencing out of GPU memory errors after running the same function repeatedly. CUDA.jl seems to be holding onto significant amounts of memory outside of the nominal memory pool that is not freed by standard GC.

Example:

julia\> CUDA.memory\_status()  
Effective GPU memory usage: 42.78% (6.376 GiB/14.903 GiB)  
CUDA allocator usage: 2.016 GiB  
Memory pool usage: 2.016 GiB (2.016 GiB allocated, 0 bytes cached)

julia\> GC.gc(true)

julia\> CUDA.memory\_status()  
Effective GPU memory usage: 39.43% (5.876 GiB/14.903 GiB)  
CUDA allocator usage: 0 bytes  
Memory pool usage: 0 bytes (0 bytes allocated, 0 bytes cached)

Checking nvidia-smi, all of this memory is indeed allocated by the julia process. I first suspected this could be a memory leak, but I have seen that this memory does appear to be occasionally freed, although I haven’t yet figured out how to regularly reproduce that. In the meantime I am getting out of memory errors in the REPL, even though all my allocations are within a single function that I am repeatedly updating and running, which presumably should not grow the overall memory footprint.

Is this expected behavior? What is this non-pool memory used for? Is there any way to force it to be freed/reclaimed? I know that device\_reset!() is broken on current CUDA distributions, but is there another way to reset and free all device memory?

UPDATE: Setting the JULIA\_CUDA\_MEMORY\_POOL environment variable to none appears to free almost all memory on GC. So perhaps I am running into some sort of fragmentation within the binned pool allocator?

with JULIA\_CUDA\_MEMORY\_POOL=none:  
julia\> CUDA.memory\_status()  
Effective GPU memory usage: 68.17% (10.159 GiB/14.903 GiB)  
CUDA allocator usage: 8.063 GiB  
Memory pool usage: 8.063 GiB (8.063 GiB allocated, 0 bytes cached)

julia\> GC.gc()

julia\> CUDA.memory\_status()  
Effective GPU memory usage: 0.64% (97.125 MiB/14.903 GiB)  
CUDA allocator usage: 0 bytes  
Memory pool usage: 0 bytes (0 bytes allocated, 0 bytes cached)

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [April 26, 2021, 10:37am UTC](https://discourse.julialang.org/t/significant-cuda-jl-memory-allocations-outside-of-main-pool/59991/2 "2021-04-26T10:37:28Z")

</div>

Yes, the binned pool has certain overheads. If you can, please upgrade to CUDA 11.2 and CUDA.jl 3.0, the new memory pool there is based on CUDA’s stream-ordered allocator, which performs better (both in terms of pooling memory, as by enabling asynchronous memory operations).

---

<div class="post-metadata">

**Author:** ![jpdoane](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jpdoane](https://discourse.julialang.org/u/jpdoane)\
**Post date:** [August 6, 2022, 7:35pm UTC](https://discourse.julialang.org/t/significant-cuda-jl-memory-allocations-outside-of-main-pool/59991/3 "2022-08-06T19:35:53Z")

</div>

I am now running CUDA.jl 3.12.0 and CUDA 11.7 but unfortunately I am continuing to wrestle with this issue, and I figured I would resurrect this old question rather than start a new thread. After running my CUDA.jl code multiple times (either from the REPL or within a loop inside a function), GPU memory seems to accumulate without being freed. This is something we’ve always fought with but were able to work around. However it has now become an issue since I am trying to perform some batch processing that is regularly aborting with OOM errors and must be manually restarted. I have tried various combinations of reclaim, gc() and device\_reset() to no avail:

```julia
julia> GC.gc(true)
julia> CUDA.reclaim()
julia> CUDA.device_reset!()
julia> CUDA.memory_status()
Effective GPU memory usage: 72.42% (34.429 GiB/47.544 GiB)
Memory pool usage: 33.151 GiB (34.156 GiB reserved)  

```

Unlike the scenario described above in my original post, memory is not simply being reserved by the pool, but seems to actually still be allocated. However, all of my the CUDA code is encapsulated behind several functions that have all returned and there are no CUDA.jl objects or data in scope as best as I can tell.

```julia
julia> varinfo()
  name size summary                                                  
  ––––––––––––––––––––––– ––––––––––– –––––––––––––––––––––––––––––––––––––––––––––––––––––––––
  Base Module                                                   
  Core Module                                                   
  InteractiveUtils 255.544 KiB Module                                                   
  Main Module                                                   
  ans 0 bytes Nothing                                                  
  batchProcess 0 bytes batchProcess (generic function with 1 method)            
  postProcess 0 bytes postProcess (generic function with 1 method)             
  sortFiles 0 bytes sortFiles (generic function with 1 method)    

```

If it matters, my code is multi-threaded, multistream, although again all of these threads have terminated. I did see [this](https://github.com/JuliaGPU/CUDA.jl/issues/866#issuecomment-827558456) which might be relevant. However, my threads that utilize CUDA.jl return “nothing”
