# AD pipeline and Hessian–vector products

**URL:** <https://discourse.julialang.org/t/ad-pipeline-and-hessian-vector-products/116329>\
**Category:** Optimization (Mathematical)\
**Tags:** autodiff\
**Created:** [June 27, 2024, 10:51pm UTC](https://discourse.julialang.org/t/ad-pipeline-and-hessian-vector-products/116329 "2024-06-27T22:51:30Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 27, 2024, 11:37pm UTC](https://discourse.julialang.org/t/ad-pipeline-and-hessian-vector-products/116329/3 "2024-06-27T23:37:23Z")

</div>

> [@RS-Coop](#):
>
> Given a twice differentiable function f:\mathbb{R}^n\to\mathbb{R}, I would like to construct an operator H\_x:\mathbb{R}^n\to\mathbb{R}^n which maps v\to\nabla^2f(x)v.

Nested AD (higher derivatives) like this can be a tricky subject, especially for large n where efficiency is important, but I think at this point there may be good solutions for the problem you are interested in.

If I understand your notation correctly, what you want is the following Hessian–vector product, which is equivalent to a directional derivative of \nabla f, and can be efficiently implemented by forward-over-reverse. i.e.

\underbrace{\frac{\partial^2 f}{\partial x^2}}\_\mathrm{Hessian} v= \left. \frac{d}{d\alpha} \left( \left. \nabla f \right|\_{x+\alpha v} \right) \right|\_{\alpha = 0}

where (for large n) the \nabla f is implemented with reverse-mode (e.g. via Enzyme.jl or Zygote.jl) and the scalar derivative d/d\alpha is implemented with forward mode (e.g. ForwardDiff.jl). In this way, you avoid ever explicitly computing the Hessian matrix, and the computational cost should be proportional to the cost of computing f(x) only once.

See e.g. the discussion and example here of a closely related problem: [Nested AD with Lux etc - #12 by stevengj](https://discourse.julialang.org/t/nested-ad-with-lux-etc/113573/12)

For example, this works fine for me:

```julia
using LinearAlgebra
f(x) = norm(x)^3 * x[end] + x[1]^2 # example ℝⁿ→ℝ function

import Zygote, ForwardDiff

# compute (∂²f/∂x²)v at x
function Hₓ(x, v)
    ∇f(y) = Zygote.gradient(f, y)[1]
    return ForwardDiff.derivative(α -> ∇f(x + α*v), 0)
end

```

I couldn’t get the analogous thing to work with Enzyme, but maybe @wsmoses knows the trick.

> [@RS-Coop](#):
>
> Further, this seems like something that should be implemented well in a single location so as to be the most reusable across the Julia ecosystem. Does this exist, and if not, where should it exist?

Yeah, it seems like DifferentiationInterface.jl should include some kind of Hessian–vector product interface. It needs to be in a higher-level package like that one because it may involve combining multiple AD packages as I did above, and the implementation can be rather non-obvious (especially because not all AD combinations support nesting).

---

_[View the full topic](https://discourse.julialang.org/t/ad-pipeline-and-hessian-vector-products/116329)._
