# Does 'gradient' allocate?

**URL:** <https://discourse.julialang.org/t/does-gradient-allocate/32880>\
**Category:** General Usage\
**Tags:** question, flux, machine-learning, zygote\
**Created:** [January 2, 2020, 5:06pm UTC](https://discourse.julialang.org/t/does-gradient-allocate/32880 "2020-01-02T17:06:47Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![e3c6](https://avatars.discourse-cdn.com/v4/letter/e/e79b87/32.png) [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Post date:** [January 2, 2020, 5:06pm UTC](https://discourse.julialang.org/t/does-gradient-allocate/32880/1 "2020-01-02T17:06:47Z")

</div>

In a call such as `g = gradient(loss, ps)` , where `ps` is a `Params` object (see for instance this line: [Flux.jl/train.jl at e92da0cf850a982c425b83c92d6274174e52b02c · FluxML/Flux.jl · GitHub](https://github.com/FluxML/Flux.jl/blob/e92da0cf850a982c425b83c92d6274174e52b02c/src/optimise/train.jl#L68)), am I allocating a new vector for the gradient? Or does the vector already exist somewhere (perhaps within the `Params` structure) and `gradient` is just returning a reference to it?

If `gradient` is allocating, would it be worthwhile to think of rewriting the `Flux.train!` loop to use a pre-allocated gradient vector?

---

<div class="post-metadata">

**Author:** ![tanhevg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tanhevg/32/12025_2.png) [@tanhevg](https://discourse.julialang.org/u/tanhevg)\
**Post date:** [January 2, 2020, 6:10pm UTC](https://discourse.julialang.org/t/does-gradient-allocate/32880/2 "2020-01-02T18:10:33Z")

</div>

The answer to the first question is yes, it will allocate. `Params` just encapsulates model parameters in various data structures (ordering vector and id set to be precise, see code for more details).

Re pre-allocating gradients - would be interesting to hear what the maintainers think of it. I would argue against it, because this would consume precious memory that could otherwise be used for temporary variables during model/gradient computation. In real life forward and backward pass takes so long, that pre-allocating a single array would not save much.
