# Zygote gradient accumulation?

**URL:** <https://discourse.julialang.org/t/zygote-gradient-accumulation/55654>\
**Category:** General Usage\
**Created:** [February 19, 2021, 9:15pm UTC](https://discourse.julialang.org/t/zygote-gradient-accumulation/55654 "2021-02-19T21:15:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jw3126](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jw3126/32/3086_2.png) [@jw3126](https://discourse.julialang.org/u/jw3126)\
**Post date:** [February 19, 2021, 9:15pm UTC](https://discourse.julialang.org/t/zygote-gradient-accumulation/55654/1 "2021-02-19T21:15:59Z")

</div>

I ported a neural network from pytorch to Flux and now I am immediately out of GPU memory.  
I would like to better understand allocations in the backward pass in Flux/Zygote.

Say I have the following toy function:

```julia
function densenet(x0)
    x1 = x0
    x2 = x0 + x1
    x3 = x0 + x1 + x2
    x4 = x0 + x1 + x2 + x3
    x5 = x0 + x1 + x2 + x3 + x4
    ...
    return xn
end

```

Now I want to compute the pullback of `densenet`. I am interested in the peak memory usage if I do this.

Mathematically, the differential x̄3 at say x3 is the sum of the differentials x̄4…x̄n.  
Is it correct, that x̄4…x̄n need to be alive in memory, for Zygote to compute x̄3?  
I think in pytorch, this is not the case and x̄3 will be accumulated incrementally as the x̄i materialize. Is that correct?

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [February 20, 2021, 12:58am UTC](https://discourse.julialang.org/t/zygote-gradient-accumulation/55654/2 "2021-02-20T00:58:32Z")

</div>

If it’s not necessary for PyTorch, then in theory it’s not necessary for Zygote either since both are performing reverse mode AD. However, there are some implementation differences that likely affect memory usage. For one, Zygote does quite a bit of copying for _certain_ operations. In-place accumulation is on the roadmap, but it’s not fully fleshed out yet. Also, PyTorch’s autograd engine is (according to [their docs](https://pytorch.org/docs/stable/notes/autograd.html#in-place-operations-with-autograd) and my observations in practice) very proactive about buffer reuse and freeing allocations early. Lastly, CUDA.jl’s allocator integrates with the Julia GC and thus isn’t guaranteed to immediately “free” (or release to a pool/arena) temporary allocations like a reference counting system would. This doesn’t matter in most cases, but can lead to OOMs when you’re close to the VRAM limit and a library like CuDNN decides to carry out its own little allocation on the side.

To help evaluate where your network allocates on the backwards pass, you could use [`Zygote.@code_adjoint`](https://fluxml.ai/Zygote.jl/latest/internals/#) or `@code_{typed/lowered/warntype}` on the pullback itself. Do post your findings in #domain:ml or GitHub—it would be great if we could catalog and prioritize some of these performance gaps.

---

<div class="post-metadata">

**Author:** ![jw3126](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jw3126/32/3086_2.png) [@jw3126](https://discourse.julialang.org/u/jw3126)\
**Post date:** [February 20, 2021, 7:15pm UTC](https://discourse.julialang.org/t/zygote-gradient-accumulation/55654/3 "2021-02-20T19:15:24Z")

</div>

Thanks for the answer! Following your suggestion I opened an issue [Incremental accumulation of gradients? · Issue #905 · FluxML/Zygote.jl · GitHub](https://github.com/FluxML/Zygote.jl/issues/905)
