# CUDA.jl with @threads causing memory leak?

**URL:** <https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853>\
**Category:** GPU\
**Tags:** bug, multithreading, cuda, threads, cudajl\
**Created:** [April 24, 2026, 7:03pm UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853 "2026-04-24T19:03:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![noetheriankoala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noetheriankoala/32/218462_2.png) [@noetheriankoala](https://discourse.julialang.org/u/noetheriankoala)\
**Post date:** [April 24, 2026, 7:03pm UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853/1 "2026-04-24T19:03:01Z")

</div>

Hello. First of all, I am using `Julia v1.12.6`, `CUDA v5.11.2`, and `FLoops v0.2.2`. Consider the following code and comments.

```julia-auto
using CUDA
using Base.Threads

function func1()
    A = CuArray{Float64}(undef, 1500, 1500, 1000)
    for i in 1:size(A,3)
        A_ = view(A, :, :, i)
    end
end

function func2()
    A = CuArray{Float64}(undef, 1500, 1500, 1000)
    @threads for i in 1:size(A,3)
        A_ = view(A, :, :, i)
    end
    return
end

# GPU almost empty
println("\nA")
CUDA.memory_status()

# do task without threading
func1()

# memory still allocated, but that is okay
println("\nB")
CUDA.memory_status()

# manually clean up
GC.gc()
CUDA.reclaim()

# GPU empty
println("\nC")
CUDA.memory_status()

# do task with threading threading
func2()

# memory still allocated, but that is okay
println("\nD")
CUDA.memory_status()

# manually clean up
GC.gc()
CUDA.reclaim()

# NOT DEALLOCATED!!!!! BUG?????
println("\nE")
CUDA.memory_status()

```

When I run this code I get the following output

```julia-auto
A
Effective GPU memory usage: 3.39% (821.250 MiB/23.643 GiB)
Memory pool usage: 0 bytes (0 bytes reserved)

B
Effective GPU memory usage: 74.37% (17.583 GiB/23.643 GiB)
Memory pool usage: 16.764 GiB (16.781 GiB reserved)

C
Effective GPU memory usage: 3.39% (821.250 MiB/23.643 GiB)
Memory pool usage: 0 bytes (0 bytes reserved)

D
Effective GPU memory usage: 74.39% (17.587 GiB/23.643 GiB)
Memory pool usage: 16.764 GiB (16.781 GiB reserved)

E
Effective GPU memory usage: 74.39% (17.587 GiB/23.643 GiB)
Memory pool usage: 16.764 GiB (16.781 GiB reserved)

```

Note that memory is not freed between `E` and `D`. What is going on?

---

<div class="post-metadata">

**Author:** ![FaultZone](https://avatars.discourse-cdn.com/v4/letter/f/e5b9ba/32.png) [@FaultZone](https://discourse.julialang.org/u/FaultZone)\
**Post date:** [April 26, 2026, 1:45am UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853/2 "2026-04-26T01:45:52Z")

</div>

Explicitly free `A`:

```julia
function func2()
    A = CuArray{Float64}(undef, 100, 100, 100)
    @threads for i in 1:size(A,3)
        A_ = view(A, :, :, i)
    end
    CUDA.unsafe_free!(A)
    return
end

```

AFAIK, it has to do with something has julia threads and the garbage collector interact.

---

<div class="post-metadata">

**Author:** ![Alexander-Barth](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alexander-barth/32/3692_2.png) [@Alexander-Barth](https://discourse.julialang.org/u/Alexander-Barth)\
**Post date:** [April 26, 2026, 9:28am UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853/3 "2026-04-26T09:28:42Z")

</div>

Does the multithreaded `func2` (without `CUDA.unsafe_free!` or `CUDA.reclaim`) actually run out-of-memory if you call it in a loop?

I guess you have seen this comment.

> [@CUDA memory isn't freed and cannot be backtracked](https://discourse.julialang.org/t/cuda-memory-isnt-freed-and-cannot-be-backtracked/83292/5):
>
> This is a perfectly normal report: you’re only using 5MB of GPU memory, while the underlying pool (which allocations are made in) is currently sized around 5GB, thus consuming most of the physical memory on your device. This does not mean that the memory is unavailable, you can allocate 5GB-5MB. So this isn’t indicative of an OOM, or a memory leak.

It would also be interesting to see what happens if you call `CUDA.reclaim` on all threads.

---

<div class="post-metadata">

**Author:** ![noetheriankoala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noetheriankoala/32/218462_2.png) [@noetheriankoala](https://discourse.julialang.org/u/noetheriankoala)\
**Post date:** [April 27, 2026, 10:23pm UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853/4 "2026-04-27T22:23:37Z")

</div>

Both the following “fix” the issue.

```julia-auto
# This is the best current work-arround
function func2()
    A = CuArray{Float64}(undef, 1500, 1500, 1000)
    @threads for i in 1:size(A,3)
        A_ = view(A, :, :, i)
        A_ = nothing
    end

    CUDA.unsafe_free!(A)

    return
end

# This also works, though I am not sure if 
# in general every thread will be utilized.
function func2()
    A = CuArray{Float64}(undef, 1500, 1500, 1000)
    @threads for i in 1:size(A,3)
        A_ = view(A, :, :, i)
        A_ = nothing
    end

    @threads for _ in 1:nthreads()
         println(threadid())
         GC.gc()
         CUDA.reclaim()
    end
    return
end

```

However, using `JULIA_CUDA_MEMORY_POOL=none` does not work. The output is as follows. In particular, note that the effective GPU memory usage is still \>17GB.

```julia-auto
A
Effective GPU memory usage: 1.66% (401.688 MiB/23.643 GiB)
No memory pool is in use.
B
Effective GPU memory usage: 72.57% (17.158 GiB/23.643 GiB)
No memory pool is in use.
C
Effective GPU memory usage: 1.66% (401.688 MiB/23.643 GiB)
No memory pool is in use.
D
Effective GPU memory usage: 72.57% (17.158 GiB/23.643 GiB)
No memory pool is in use.
E
Effective GPU memory usage: 72.57% (17.158 GiB/23.643 GiB)

```

This looks like a bug to me. Do you agree, and if so, where do you suggest I report it?
