# Higher order derivatives in Flux

**URL:** <https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051>\
**Category:** Machine Learning\
**Tags:** differentiation\
**Created:** [May 21, 2018, 4:28pm UTC](https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051 "2018-05-21T16:28:22Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [May 21, 2018, 4:28pm UTC](https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051/1 "2018-05-21T16:28:22Z")

</div>

As I understand it, Flux.jl does not support the calculation of higher-order derivatives as of yet (please correct me if I’m wrong). Are there any plans to support this in the future? I opened an issue a while back [Flux.jl/issues/129](https://github.com/FluxML/Flux.jl/issues/129) but it hasn’t generated any discussion. My use cases are two:

- Calculating Hessian-vector products [see the issue for details](https://github.com/FluxML/Flux.jl/issues/129)
- Including gradient terms in a cost function (a regularization application)

If HOD are not on the Flux horizon, I would very much appreciate suggestions on manual workarounds (Reverse/ForwardDiff etc.) that could integrate with an existing Flux-based workflow.

Thanks!

---

<div class="post-metadata">

**Author:** ![crinders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/crinders/32/8434_2.png) [@crinders](https://discourse.julialang.org/u/crinders)\
**Post date:** [May 22, 2018, 12:28pm UTC](https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051/2 "2018-05-22T12:28:22Z")

</div>

If all you need is the Hessian, then `ReverseDiff.hessian(f,x)` should do what you want. If you want higher derivatives, repeated application of `ReverseDiff.jacobian` will do the trick. For instance,

```julia
f = x-> sum(exp.(x))^2
g = x->ReverseDiff.gradient(f,x)
h = x->ReverseDiff.jacobian(g,x)
ReverseDiff.jacbobian(h,[1.0,-3.0]

```

The last line will give you the third order derivative tensor of `f` for `[1.0,-3.0]` as a matrix, i.e., in order to get the tensor, you would have to change the shape of the output.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [May 22, 2018, 12:49pm UTC](https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051/3 "2018-05-22T12:49:15Z")

</div>

Thanks, those methods are useful. What I’m concerned about is how to integrate them with the training loop and optimizers written in Flux. I see no straightforward way of making use of ReverseDiff hessians with the Flux optimizers 😕

---

<div class="post-metadata">

**Author:** ![crinders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/crinders/32/8434_2.png) [@crinders](https://discourse.julialang.org/u/crinders)\
**Post date:** [May 22, 2018, 3:08pm UTC](https://discourse.julialang.org/t/higher-order-derivatives-in-flux/11051/4 "2018-05-22T15:08:56Z")

</div>

I’ll let the Flux.jl people answer, but worst-case, you will have to pass the ReverseDiff results to the tracked variable’s gradient data field, e.g.,

```julia
X.grad[:]=ReverseDiff.hessian(f,X.data)*X.data +ReverseDiff.gradient(f,X.data)

```

if your objective is (\nabla^T f) x. I don’t know whether that has any nasty side-effects.
