# Is there an efficient way to compute the Hessian of a NN?

**URL:** <https://discourse.julialang.org/t/is-there-an-efficient-way-to-compute-the-hessian-of-a-nn/26971>\
**Category:** Machine Learning\
**Tags:** flux\
**Created:** [July 30, 2019, 10:08am UTC](https://discourse.julialang.org/t/is-there-an-efficient-way-to-compute-the-hessian-of-a-nn/26971 "2019-07-30T10:08:33Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![HenriDeh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrideh/32/8316_2.png) [@HenriDeh](https://discourse.julialang.org/u/HenriDeh)\
**Post date:** [August 2, 2019, 2:44pm UTC](https://discourse.julialang.org/t/is-there-an-efficient-way-to-compute-the-hessian-of-a-nn/26971/4 "2019-08-02T14:44:52Z")

</div>

Thank you both for your answers. And thank you for this nice paper, I understand AD much better now. I wanted to attempt both RoF and double forward methods. But I encounter a stack overflow issue when I use ForwardDiff methods on a neural network. Only the jacobian method takes an `AbstractArray` as input and accepts `AbstractArrays` as output. The other assume one of the two is Real. However:

```julia
f = Chain(Dense(10,5,elu),Dense(5,3))
x = rand(10)
ForwardDiff.jacobian(f, x) 

```

Throws a stack overflow. I assume this method does n forward passes, one for each input. But the underlying `ForwardDiff.gradient` cannot take multi-dimensional output functions as argument. Isn’t it weird for a forward accumulation mode to not be able to do that ? According to the paper from Baydin et al. it is quite straightforward to do. Same for the `ForwardDiff.derivative` function, it only handle functions with `Real ` inputs.

Am I missing something ?

---

_[View the full topic](https://discourse.julialang.org/t/is-there-an-efficient-way-to-compute-the-hessian-of-a-nn/26971)._
