# ReverseDiff.GradientTapes for functions of the form x -\> f(x, y) when gradient is to be calculated wrt x only?

**URL:** <https://discourse.julialang.org/t/reversediff-gradienttapes-for-functions-of-the-form-x-f-x-y-when-gradient-is-to-be-calculated-wrt-x-only/30668>\
**Category:** General Usage\
**Created:** [November 3, 2019, 10:27pm UTC](https://discourse.julialang.org/t/reversediff-gradienttapes-for-functions-of-the-form-x-f-x-y-when-gradient-is-to-be-calculated-wrt-x-only/30668 "2019-11-03T22:27:32Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![bashonubuntu](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bashonubuntu](https://discourse.julialang.org/u/bashonubuntu)\
**Post date:** [November 3, 2019, 10:27pm UTC](https://discourse.julialang.org/t/reversediff-gradienttapes-for-functions-of-the-form-x-f-x-y-when-gradient-is-to-be-calculated-wrt-x-only/30668/1 "2019-11-03T22:27:32Z")

</div>

Hello,

I typically differentiate log-likelihood functions of the forms x → f(x, data) wrt x, where x is a multi-dimensional input, data is some _fixed_ dataframe, f is a scalar-valued function. To do so, I typically use either ForwardDiff or ReverseDiff. I was recently looking back at the ReverseDiff documentation and there was a recommendation to use a ReverseDiff.GradientTape to prerecord f. I was wondering if there is way to pre-complie “tapes” for functions of the form x → f(x, data). The relevant links which made me ask this are

[http://www.juliadiff.org/ReverseDiff.jl/api/#the-abstracttape-api](http://www.juliadiff.org/ReverseDiff.jl/api/#the-abstracttape-api)

[https://github.com/JuliaDiff/ReverseDiff.jl/blob/master/examples/gradient.jl](https://github.com/JuliaDiff/ReverseDiff.jl/blob/master/examples/gradient.jl)

While searching for a solution, I have also read about talks of Capstan.jl but I am unable to appreciate how this new package (and, there seems to be a lot of excitement about this!) will improve the existing implementations of ForwardDiff.jl and ReverseDiff.jl.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [November 4, 2019, 6:50am UTC](https://discourse.julialang.org/t/reversediff-gradienttapes-for-functions-of-the-form-x-f-x-y-when-gradient-is-to-be-calculated-wrt-x-only/30668/2 "2019-11-04T06:50:10Z")

</div>

I found tapes to be very fragile in practice (a lot of seemingly innocuous Julia code has branches, which will break things).

I have created a simple interface package

> **[GitHub - tpapp/LogDensityProblems.jl: A common framework for implementing and...](https://github.com/tpapp/LogDensityProblems.jl)**
>
> A common framework for implementing and using log densities for inference. - GitHub - tpapp/LogDensityProblems.jl: A common framework for implementing and using log densities for inference.

which allows you to define a \mathbb{R}^n \to \mathbb{R} callable (so put the data in eg a `struct`), then AD it via either ForwardDiff, ReverseDiff, Flux, or Zygote (experimental).

The [docs](https://tamaspapp.eu/LogDensityProblems.jl/dev/) has a worked example.

---

<div class="post-metadata">

**Author:** ![bashonubuntu](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bashonubuntu](https://discourse.julialang.org/u/bashonubuntu)\
**Post date:** [November 4, 2019, 2:20pm UTC](https://discourse.julialang.org/t/reversediff-gradienttapes-for-functions-of-the-form-x-f-x-y-when-gradient-is-to-be-calculated-wrt-x-only/30668/3 "2019-11-04T14:20:20Z")

</div>

Thanks, this seems very helpful and I will check it out.

Yeah, my experience with ReverseDiff has been the same. I have found ForwardDiff to be the most robust. Has this been anyone else’s experience?

My go to AD code for say a function from \mathbb{R}^2 to \mathbb{R} looks something like

```julia
function calculate_gradient(x, data, cfg)
 ForwardDiff.gradient!(Array{Float64}(undef, 1, 2), x -> f(x, data), beta, cfg, Val{false}())
end

```

where

```julia
const cfg = ForwardDiff.GradientConfig(x -> f(x, y), beta, ForwardDiff.Chunk{2}())

```

and I have found this to be the most robust in my experiments, even though the documentation suggests that one should use ReverseDiff for problems where f:\mathbb{R}^n \rightarrow \mathbb{R}, where n \> 1 (but also mentions that ForwardDiff can be faster for low dimensional inputs).

[https://github.com/JuliaDiff/ReverseDiff.jl](https://github.com/JuliaDiff/ReverseDiff.jl)
