# Is garbage collection bad for gpu programming?

**URL:** <https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942>\
**Category:** GPU\
**Tags:** garbage-collection\
**Created:** [October 24, 2020, 4:04pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942 "2020-10-24T16:04:59Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 24, 2020, 4:04pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/1 "2020-10-24T16:04:59Z")

</div>

I was hearing this podcast:

> **[The Great ML Language (Un)Debate](https://twimlai.com/article/the-great-ml-language-undebate/)**
>
> In this panel, we bring together experts and enthusiasts representing an array of both popular and emerging programming languages for machine learning.

and the Swift guy at a moment (about minute 56) said something like that no language would replace Python if one of the qualities of Python is not followed, concerning not having garbage collection, because that was important for performant GPU programming. The phrase seemed directed to Julia in the context.

I am completely ignorant about these themes, but is there any reason to believe that Julia is not suitable for high-performance GPU computing in any sense?

---

<div class="post-metadata">

**Author:** ![DoktorMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/doktormike/32/2736_2.png) [@DoktorMike](https://discourse.julialang.org/u/DoktorMike)\
**Post date:** [October 24, 2020, 5:32pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/2 "2020-10-24T17:32:49Z")

</div>

Not sure I understand his point. Both python and julia have garbage collection. They just go about it in different ways. I haven’t read the post so I cannot really say more.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 24, 2020, 6:11pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/3 "2020-10-24T18:11:52Z")

</div>

The part of this that is true is that garbage collection on GPU really sucks. The Julia solution to that is to write code that only creates objects on the stack so GC isn’t needed.

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [October 24, 2020, 7:35pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/4 "2020-10-24T19:35:07Z")

</div>

The main distinction between Python and Julia here is that many people feel that Python’s reference counting mechanism makes it bit easier to regulate when GPU memory is freed than Julia’s GC.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 24, 2020, 10:48pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/5 "2020-10-24T22:48:55Z")

</div>

> [@Oscar\_Smith](#):
>
> creates objects on the stack

Does this mean something very specific in Julia? (As using static arrays, or tuples, etc?)

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 24, 2020, 10:56pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/6 "2020-10-24T22:56:45Z")

</div>

The concept isn’t Julia specific (C also has something similar where malloc allocates to the heap, everything else goes to the stack). Julia currently stack allocates any immutable struct (or primitive type). That said, this behavior is 1.5 and newer. [Julia 1.5 Highlights](https://julialang.org/blog/2020/08/julia-1.5-highlights/#struct_layout_and_allocation_optimizations) gives a very good analysis of the change.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 25, 2020, 12:56am UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/7 "2020-10-25T00:56:26Z")

</div>

That is an excellent article. I am curious about one statement:

“As a result of this work, arbitrary immutable objects—regardless of whether they have fields that reference mutable objects or not—can now be stack allocated”

Does that mean that an array that is part of an immutable struct is treated differently than an array not within a struct? If so, can this difference be important for performance? (Would it make sense to declare an struct containing a single array just because of that?)

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 25, 2020, 12:59am UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/8 "2020-10-25T00:59:23Z")

</div>

No. When a struct `T` has an `Array` as a field, it is storing a pointer to that array. Prior to 1.5, this would mean that `T` had to be heap allocated. In 1.5, the `Array` is still heap allocated, but the struct is stack allocated. There is no difference for the `Array`, just for things that reference it.

---

<div class="post-metadata">

**Author:** ![jonathan-laurent](https://avatars.discourse-cdn.com/v4/letter/j/ecae2f/32.png) [@jonathan-laurent](https://discourse.julialang.org/u/jonathan-laurent)\
**Post date:** [October 25, 2020, 1:40am UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/9 "2020-10-25T01:40:40Z")

</div>

I would be interested in hearing more about this. Is this a practical solution in most cases? Isn’t Flux still allocating a whole lot for example?

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [October 25, 2020, 2:46pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/10 "2020-10-25T14:46:48Z")

</div>

> [@Oscar\_Smith](#):
>
> Prior to 1.5, this would mean that `T` had to be heap allocated. In 1.5, the `Array` is still heap allocated, but the struct is stack allocated.

No on both versions the struct can be either heap or stack allocated. On 1.5 the struct will be inlined.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [October 25, 2020, 3:58pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/11 "2020-10-25T15:58:09Z")

</div>

On the other hand, it’s much harder to write Python code that doesn’t allocate in the first place, whereas Julia has lots of non-allocating in-place APIs.

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [October 25, 2020, 4:23pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/12 "2020-10-25T16:23:59Z")

</div>

True, but I don’t think that matters as much for something like PyTorch, which I don’t think aspires to be a general purpose GPU language.

---

<div class="post-metadata">

**Author:** ![Azamat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/azamat/32/6892_2.png) [@Azamat](https://discourse.julialang.org/u/Azamat)\
**Post date:** [October 25, 2020, 7:43pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/13 "2020-10-25T19:43:04Z")

</div>

It would be really helpful for this discussion if someone can outline all the differences between the approaches to garbage collection in Julia and Python and what are the tradeoffs, advantages, and disadvantages in each (including with regards to GPU)

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 25, 2020, 8:05pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/14 "2020-10-25T20:05:00Z")

</div>

These texts seem nice:

> **[Master’s Thesis: Garbage Collection Abstractions for High-Level GPU Languages](https://jonathanvdc.github.io/portfolio/2019-thesis/)**
>
> I created the first GPU garbage collector for the Julia language!

[Garbage Collection (.NET) vs. ARC (Swift) - Part 2](https://www.jbssolutions.com/resources/blog/garbage-collection-net-vs-arc-swift/#:~:text=Automatic%20Reference%20Counting%20is%20technically,part%20of%20your%20application%20code).

---

<div class="post-metadata">

**Author:** ![Akatz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/akatz/32/15164_2.png) [@Akatz](https://discourse.julialang.org/u/Akatz)\
**Post date:** [October 25, 2020, 9:05pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/15 "2020-10-25T21:05:14Z")

</div>

See this [Julia github issue](https://github.com/JuliaLang/julia/issues/34836) for a related discussion

According to Jeff and Keno, compile time memory management seems to be a promising alternative.

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [October 25, 2020, 11:34pm UTC](https://discourse.julialang.org/t/is-garbage-collection-bad-for-gpu-programming/48942/16 "2020-10-25T23:34:48Z")

</div>

Unfortunately, the implementation of this in CUDAnative.jl was not merged because (from what I’ve heard), the implementation relied on a GPU-to-CPU communication mechanism that was not performant enough for general usage.

Continual memory allocation/deallocation is the thing which causes the GC to work hard, so it’s the thing that needs to be tackled first. I think what will be necessary is for Flux to perform more efficient memory management, specifically trying to reuse previously-allocated buffers when possible. To my knowledge, Flux doesn’t worry about these things in the interest of simplicity (I could of course be wrong).
