# Memory usage seems extremely high in standard Flux, is it any better with Zygote?

**URL:** <https://discourse.julialang.org/t/memory-usage-seems-extremely-high-in-standard-flux-is-it-any-better-with-zygote/27435>\
**Category:** Machine Learning\
**Created:** [August 11, 2019, 11:29pm UTC](https://discourse.julialang.org/t/memory-usage-seems-extremely-high-in-standard-flux-is-it-any-better-with-zygote/27435 "2019-08-11T23:29:34Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![tue](https://avatars.discourse-cdn.com/v4/letter/t/4491bb/32.png) [@tue](https://discourse.julialang.org/u/tue)\
**Post date:** [August 11, 2019, 11:29pm UTC](https://discourse.julialang.org/t/memory-usage-seems-extremely-high-in-standard-flux-is-it-any-better-with-zygote/27435/1 "2019-08-11T23:29:34Z")

</div>

Hi I’m running some tests on CIFAR10, with a residual CNN, and so far I’m getting very discouraging results for the memory consumption. I’m doing something rather atypical where I regularize each layer of my network separately with a graph Laplacian, however it seems to take up way more memory than I would have imagined it should based on what I have seen previously in pytorch.  
with a batch size of 100 it quickly eats up all 32GB of memory before I even reach the backpropagation!

For comparison:  
I run with 1.5 M parameters in Pytorch, and a batch size of 200, and it consumes about 8 GB of memory.  
In Flux I had to downscale severly and still I’m having trouble… I run with the last couple of layers cut off (otherwise same network), so I have 350k parameters. I run a batch size of 10 images, and it consumes ~30 GB of memory.

Should I expect the same behaviour with Zygote?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [August 12, 2019, 12:56am UTC](https://discourse.julialang.org/t/memory-usage-seems-extremely-high-in-standard-flux-is-it-any-better-with-zygote/27435/2 "2019-08-12T00:56:27Z")

</div>

Not directly answering the question and I’m not sure if it is related, but for reference:

[https://github.com/JuliaGPU/CuArrays.jl/issues/323](https://github.com/JuliaGPU/CuArrays.jl/issues/323)

---

<div class="post-metadata">

**Author:** ![tue](https://avatars.discourse-cdn.com/v4/letter/t/4491bb/32.png) [@tue](https://discourse.julialang.org/u/tue)\
**Post date:** [August 12, 2019, 7:56am UTC](https://discourse.julialang.org/t/memory-usage-seems-extremely-high-in-standard-flux-is-it-any-better-with-zygote/27435/3 "2019-08-12T07:56:20Z")

</div>

Mine happens even in the first iteration, and furthermore I don’t use GPU/CuArrays at all at this point because of the large memory consumption, so I doubt it is related. But I’m guessing there is a difference in how the autograd is made between Pytorch and Flux?
