# Using the gradient of wrt the input in the loss function

**URL:** <https://discourse.julialang.org/t/using-the-gradient-of-wrt-the-input-in-the-loss-function/114983>\
**Category:** Machine Learning\
**Tags:** question\
**Created:** [May 30, 2024, 8:08pm UTC](https://discourse.julialang.org/t/using-the-gradient-of-wrt-the-input-in-the-loss-function/114983 "2024-05-30T20:08:11Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [May 30, 2024, 9:58pm UTC](https://discourse.julialang.org/t/using-the-gradient-of-wrt-the-input-in-the-loss-function/114983/2 "2024-05-30T21:58:19Z")

</div>

I think this is related, ping @avikpal

> [@Nested AD with Lux etc](https://discourse.julialang.org/t/nested-ad-with-lux-etc/113573):
>
> Nested AD (Starting v0.5.38) Starting v0.5.38, Lux automatically captures some common possibilities of AD calls inside loss functions on lux layers and converts them into a faster version to do a JVP over gradient (instead of the default reverse over reverse). In short, this finally handles several complaints we have had over the years of not being able to handle nested AD of Neural Networks efficiently. See [Nested Automatic Differentiation | Lux.jl Documentation](https://lux.csail.mit.edu/stable/manual/nested_autodiff) function loss\_function2(model, …

---

_[View the full topic](https://discourse.julialang.org/t/using-the-gradient-of-wrt-the-input-in-the-loss-function/114983)._
