# \[ANN\] Yota.jl - yet another reverse-mode autodiff package

**URL:** <https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139>\
**Category:** Package Announcements\
**Created:** [February 24, 2019, 12:11pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139 "2019-02-24T12:11:09Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 24, 2019, 12:11pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/1 "2019-02-24T12:11:09Z")

</div>

[**Yota.jl**](https://github.com/dfdx/Yota.jl) is a package for **reverse-mode automatic differentiation** designed specifically for machine learning applications.

* * *

### Usage

```julia
mutable struct Linear{T}
    W::AbstractArray{T,2}
    b::AbstractArray{T}
end

forward(m::Linear, X) = m.W * X

loss(m::Linear, X) = sum(forward(m, X))

m = Linear(rand(3,4), rand(3))
X = rand(4,5)

val, g = grad(loss, m, X)

```

where `g` is an object of type `GradientResult` holding gradients w.r.t. input variables. For scalars and tensors it returns gradient value, for structs it returns dictionary of (field path → gradient) pairs:

```julia-auto
julia> g[1]
Dict{Tuple{Symbol},Array{Float64,2}} with 1 entry:
  (:W,) => [3.38128 2.97142 2.39706 1.55525; 3.38128 2.97142 2.39706 1.55525; 3.38128 2.97142 2.39706 1.55525] # gradient w.r.t. m.W

julia> g[2] # gradient w.r.t. X
4×5 Array{Float64,2}:
 0.910691 0.910691 0.910691 0.910691 0.910691
 1.64994 1.64994 1.64994 1.64994 1.64994
 1.81215 1.81215 1.81215 1.81215 1.81215
 2.31594 2.31594 2.31594 2.31594 2.31594

```

`GradientResult` can be used in conjunction with `update!()` function to modify tensors and fields of (mutable) structs, see [README](https://github.com/dfdx/Yota.jl/blob/master/README.md) for more details.

* * *

### Features

- differentiation over **scalars** , **tensors** and **structs** designed to support PyTorch-like API
- easy to add **custom derivatives**
- [**Cassette**](https://github.com/jrevels/Cassette.jl/)-based tracer which avoids structural type constraints
- experimental GPU support via CuArrays
- performance-first implementation

Note that tracer is fully customizable and can be used independently of automatic differentiation. Again, see [README](https://github.com/dfdx/Yota.jl/blob/master/README.md#tracer-and-the-tape) for an example.

### Performance

Comparison of autodiff implementations turns out to be unexpectedly hard because of different sets of supported features, different sets of primitives, etc. (e.g. see an [attempt](https://github.com/FluxML/Zygote.jl/issues/67) to compare it with Zygote.jl). However, Yota uses a number of proven optimizations from my previous autodiff packages, so generally performance of differentiation pass should be not larger than 2-3x compared to a call to the original function. To put it differently, if you see significant difference between automatic and manually-crafted differentiation, **consider it a bug**.

### Comparison to other autodiff packages

Unlike [**Zygote**](https://github.com/FluxML/Zygote.jl/), Yota emphasizes performance over flexiblity. While Zygote aims to support full dynamism of the Julia language, Yota restricts a user to a static graph consisting of analytical functions commonly used in ML. This way we can apply a number of optimizations including memory buffer pre-allocation, common subexpression elimination, etc.

Unlike [**Capstan**](https://github.com/JuliaDiff/Capstan.jl), Yota is implemented 🙂 Although I’ll be very curious to test Capstan when it’s shipped.

Also Yota doesn’t use tracked arrays or function overloading like [**AutoGrad**](https://github.com/denizyuret/AutoGrad.jl), [**ReverseDiff**](https://github.com/JuliaDiff/ReverseDiff.jl) or current [**Flux**](https://github.com/FluxML/Flux.jl) tracker and thus doesn’t hit ambiguity issues of multiple dispatch. The downside is that dynamic computational graphs become harder to implement and are currently not supported in Yota. For the curious, there’s also an older version of the package stored in [YotaTracked](https://github.com/dfdx/YotaTracked.jl) and implemented similar to the mentioned libraries.

Feedback and bug reports are welcome!

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [February 24, 2019, 2:15pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/2 "2019-02-24T14:15:18Z")

</div>

Thanks this looks interesting.

I see that you separate the definitions for multiple arguments, which I presume means that each is only evaluated if required:

```julia
@diffrule ^(x::Real, y::Real) x y * x ^ (y-1) * ds
@diffrule ^(x::Real, y::Real) y log(x) * x ^ y * ds

```

Is there any way to share information, e.g. here to re-use `x ^ y` from the forward pass? Or perhaps this falls under common subexpression elimination & happens automatically?

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [February 24, 2019, 2:27pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/3 "2019-02-24T14:27:26Z")

</div>

Looks very interesting! Does Yota.jl currently support nested and higher-order differentiation?

---

<div class="post-metadata">

**Author:** ![saschatimme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saschatimme/32/10313_2.png) [@saschatimme](https://discourse.julialang.org/u/saschatimme)\
**Post date:** [February 24, 2019, 3:42pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/4 "2019-02-24T15:42:15Z")

</div>

Very cool! Does Yota support complex differentiation?

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [February 24, 2019, 6:27pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/5 "2019-02-24T18:27:54Z")

</div>

A related comment: is there a mathematical text that goes over what automatic differentiation is?

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 24, 2019, 6:38pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/6 "2019-02-24T18:38:18Z")

</div>

> [@improbable22](#):
>
> Is there any way to share information, e.g. here to re-use `x ^ y` from the forward pass? Or perhaps this falls under common subexpression elimination & happens automatically?

Exactly, CSE should take care of it. Right now it’s done during function compilation and is somewhat hard to observe, in future I’m planning to make it a part of tape transformation so one would be able to see exact instructions to be executed.

Also note that if one of parameters is constant (e.g. in expression `x ^ 2.0`), derivative w.r.t. this param isn’t recorded at all.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 24, 2019, 7:03pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/7 "2019-02-24T19:03:10Z")

</div>

Not yet, but it would be interesting feature to add! However, if you think about interface like:

```julia
f(x) = ...
g(x) = grad(f, x)[2] * ...

```

most likely it won’t work, at least not with the current implementation of `grad` - currently `grad` works as a caching layer for underlying `_grad` which itself returns a compiled tape with pre-allocated buffers. So it’s not like taking one ordinary function and returning another ordinary function. But with appropriate API I believe we can make higher order derivatives on the tape level and still have all the optimizations applied to the final result.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [February 24, 2019, 7:08pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/8 "2019-02-24T19:08:58Z")

</div>

> [@affans](#):
>
> A related comment: is there a mathematical text that goes over what automatic differentiation is?

There are quite a few good texts, I think

```bibtex
@book{griewank2008evaluating,
  title = {Evaluating derivatives: principles and techniques of algorithmic differentiation},
  author = {Griewank, Andreas and Walther, Andrea},
  volume = {105},
  year = {2008},
  publisher = {Siam},
}

```

is a good introduction.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 24, 2019, 7:10pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/9 "2019-02-24T19:10:33Z")

</div>

> [@saschatimme](#):
>
> Very cool! Does Yota support complex differentiation?

Haven’t tried it, but Yota should be general enough to support any kind of differentiation that boils down to primitives and chain rule. Do you have an example of you need at hand?

> [@affans](#):
>
> A related comment: is there a mathematical text that goes over what automatic differentiation is?

Not really a mathematical text, but I once explained reverse-mode autodiff example [here](https://stats.stackexchange.com/questions/224140/step-by-step-example-of-reverse-mode-automatic-differentiation/235758#235758).

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [February 24, 2019, 7:13pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/10 "2019-02-24T19:13:19Z")

</div>

> [@dfdx](#):
>
> CSE should take care of it.

Awesome. But for more complicated functions, say I have something where I’d like to cache some large calculation from which both forward & backward are then fast. Can I saftely jut place this in `mybigfunc(x)` and call that twice (with identical arguments)?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 24, 2019, 7:34pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/11 "2019-02-24T19:34:06Z")

</div>

CSE can fail on even simple functions, like `exp`.

```julia
julia> rx = Ref(0.5);

julia> using BenchmarkTools

julia> foo(x) = exp(x)
foo (generic function with 1 method)

julia> ∂foo(x) = (exp(x),exp(x))
∂foo (generic function with 1 method)

julia> @benchmark foo($(rx)[])
BenchmarkTools.Trial: 
  memory estimate: 0 bytes
  allocs estimate: 0
  --------------
  minimum time: 6.350 ns (0.00% GC)
  median time: 6.450 ns (0.00% GC)
  mean time: 6.554 ns (0.00% GC)
  maximum time: 21.960 ns (0.00% GC)
  --------------
  samples: 10000
  evals/sample: 1000

julia> @benchmark ∂foo($(rx)[])
BenchmarkTools.Trial: 
  memory estimate: 0 bytes
  allocs estimate: 0
  --------------
  minimum time: 13.317 ns (0.00% GC)
  median time: 13.778 ns (0.00% GC)
  mean time: 14.066 ns (0.00% GC)
  maximum time: 28.647 ns (0.00% GC)
  --------------
  samples: 10000
  evals/sample: 998

```

Which means, unless I’m mistaken, `logistic(x)` in this example is going to get computed three times:

```julia
logistic(x) = 1 / (1 + exp(-x))
# for an expression like `logistic(x)` where x is a Number
# gradient w.r.t. x
# is `(logistic(x) * (1 - logistic(x)) * ds)` where "ds" stands for derivative "dL/dy"
@diffrule logistic(x::Number) x (logistic(x) * (1 - logistic(x)) * ds)

```

Once on the forward pass, and twice on the backwards pass.

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [February 24, 2019, 7:43pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/12 "2019-02-24T19:43:22Z")

</div>

OK, so the CSE in question here is just Julia’s? I had the idea that the macro `@diffrule` was itself looking for such repeats… but perhaps that is the planned tape feature mentioned?

Maybe I can answer this by benchmarking gradients of `logistic(x)` when I have a minute.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 24, 2019, 8:05pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/13 "2019-02-24T20:05:50Z")

</div>

With libraries like [CommonSubexpressions](https://github.com/rdeits/CommonSubexpressions.jl) or [DataFlow](https://github.com/MikeInnes/DataFlow.jl), it’s certainly possible to do it on the Julia side!

Maybe Yota does. I haven’t tested that / dived into the source code.  
The hard part is you need some form (weaker than Julia’s `@pure`) of purity, to know functions like `logistic` are going to be the same each time they’re called.

Given that Yota emphasizes performance, and already requires static graphs, I could certainly see the case for some side-effect free assumptions to allow aggressive CSE.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 24, 2019, 8:33pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/14 "2019-02-24T20:33:46Z")

</div>

> [@Elrod](#):
>
> Given that Yota emphasizes performance, and already requires static graphs, I could certainly see the case for some side-effect free assumptions to allow aggressive CSE.

And this is exactly how it works 🙂 To be precise, here’s [the line](https://github.com/dfdx/Yota.jl/blob/master/src/compile.jl#L86) that performs CSE. You can check the result like this (which is not the part of external API though):

```julia
julia> using Yota

julia> foo(x) = exp(x)
foo (generic function with 1 method)

julia> _, g = grad(foo, 0.5)
(1.6487212707001282, GradResult(1))

julia> g.tape
Tape
  inp %1::Float64
  %2 = exp(%1)::Float64
  const %3 = 1.0::Float32
  %4 = exp(%1)::Float64
  %5 = *(%4, %3)::Float64

julia> Yota.generate_function_expr(g.tape)
:(function ##tape_fn#364()
      #= /home/slipslop/work/Yota/src/compile.jl:112 =#
      #= prologue:0 =#
      %1 = (inp %1::Float64).val
      %2 = (%2 = exp(%1)::Float64).val
      %3 = (const %3 = 1.0::Float32).val
      %4 = (%4 = exp(%1)::Float64).val
      %5 = (%5 = *(%4, %3)::Float64).val
      #= body:0 =#
      #= /home/slipslop/.julia/packages/Espresso/Lewh0/src/exgraph.jl:100 =#
      %2 = (exp)(%1)
      %3 = 1.0f0
      %5 = (*)(%2, %3)
      #= epilogue:0 =#
      (inp %1::Float64).val = %1
      (%2 = exp(%1)::Float64).val = %2
      (const %3 = 1.0::Float32).val = %3
      (%4 = exp(%1)::Float64).val = %4
      (%5 = *(%4, %3)::Float64).val = %5
  end)

```

Except for prologue and epilogue (which are used for buffer pre-allocation and are the most important optimization for large-scale deep learning models), the only code left is:

```julia
%2 = (exp)(%1)
%3 = 1.0f0
%5 = (*)(%2, %3)

```

We could further optimize it to just `exp(%1)`, but usually it doesn’t make much difference and is presumably eliminated by Julia compiler anyway.

It might look limiting to forbid mutable operations, but for comparison here’s [citation of PyTorch documentation](https://pytorch.org/docs/stable/notes/autograd.html#in-place-operations-with-autograd):

> Supporting in-place operations in autograd is a hard matter, and we discourage their use in most cases. Autograd’s aggressive buffer freeing and reuse makes it very efficient and there are very few occasions when in-place operations actually lower memory usage by any significant amount. Unless you’re operating under heavy memory pressure, you might never need to use them.

---

<div class="post-metadata">

**Author:** ![saschatimme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saschatimme/32/10313_2.png) [@saschatimme](https://discourse.julialang.org/u/saschatimme)\
**Post date:** [February 24, 2019, 8:36pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/15 "2019-02-24T20:36:04Z")

</div>

> Haven’t tried it, but Yota should be general enough to support any kind of differentiation that boils down to primitives and chain rule. Do you have an example of you need at hand?

I need derivatives of multivariate polynomials with real or complex coefficients and maybe some basic analytic functions (exp, log, sin, cos). So far I did not find an AD package which doesn’t introduce significant overhead.

An example would be the derivative of

```julia
 f(x) = x[3]*x[8]^3 - 3*x[3]*x[8]*x[6]^2 - x[1]*x[6]^3 + 3*x[1]*x[6]*x[8]^2 + x[4]*x[5]^3 - 3*x[4]*x[5]*x[7]^2 - x[2]*x[7]^3 + 3*x[2]*x[7]*x[5]^2 - 1.2342523

```

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [February 24, 2019, 11:04pm UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/16 "2019-02-24T23:04:00Z")

</div>

Do you want symbolic derivatives? It sounds like you just need a polynomial manipulator.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [February 25, 2019, 6:15am UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/17 "2019-02-25T06:15:22Z")

</div>

This is probably more general question.

I wonder if it would not be possible to share a differential rules for libraries. There are [GitHub - JuliaDiff/DiffRules.jl: A simple shared suite of common derivative definitions](https://github.com/JuliaDiff/DiffRules.jl) and I wonder if Yota could not use those. Since I work mostly with static graphs, I would be curious to try Yota. On the other hand I use few custom gradients, which I have currently made available for Flux.

But I am already impressed by Yota.

---

<div class="post-metadata">

**Author:** ![saschatimme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saschatimme/32/10313_2.png) [@saschatimme](https://discourse.julialang.org/u/saschatimme)\
**Post date:** [February 25, 2019, 6:59am UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/18 "2019-02-25T06:59:27Z")

</div>

Sure that is possible, but there are a lot of applications where you have your polynomials already in a factored form and the expanded symbolic form would be quite wasteful.

For example take a linear form `l(x) = x[1] + 3x[2]+ 2x[3] + 5x[4]` and the polynomial  
`g(x) = l(x)^4 - l(x)^2`. You really want to avoid to compute with the symbolic form of this.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 25, 2019, 7:07am UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/19 "2019-02-25T07:07:37Z")

</div>

> [@improbable22](#):
>
> Awesome. But for more complicated functions, say I have something where I’d like to cache some large calculation from which both forward & backward are then fast. Can I saftely jut place this in `mybigfunc(x)` and call that twice (with identical arguments)?

Usually yes, but as with any early-stage software, it’s worth to check it in practice 🙂 Right now you can use code from [this reply](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/14) to inspect generated stuff, I also created an [issue](https://github.com/dfdx/Yota.jl/issues/22) to make it more transparent.

> [@saschatimme](#):
>
> An example would be the derivative of

One possible issue is that you use lots of `x[i]` expressions which themselves have derivatives like:

```julia
dx = zero(x)
dx[i] .= dy

```

So in your example you will get 20 intermediate arrays of the same size as `x` and most libraries - including Yota at its current state - won’t optimize them out. Let me see what we can do in this case (although I can’t promise great results since Yota is mostly optimized for quite large arrays).

Also thanks for this example, it helped to uncover a bug in derivatives of multiargument \*, /, etc. I’ve created an [issue](https://github.com/dfdx/Yota.jl/issues/23) for this.

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [February 25, 2019, 8:26am UTC](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139/20 "2019-02-25T08:26:07Z")

</div>

> [@dfdx](#):
>
> It might look limiting to forbid mutable operations, but for comparison here’s [citation of PyTorch documentation](https://pytorch.org/docs/stable/notes/autograd.html#in-place-operations-with-autograd):
> 
> > Supporting in-place operations in autograd is a hard matter, and we discourage their use in most cases. Autograd’s aggressive buffer freeing and reuse makes it very efficient and there are very few occasions when in-place operations actually lower memory usage by any significant amount. Unless you’re operating under heavy memory pressure, you might never need to use them.

The problem with this is that it is at odds with recommendations for performance in general Julia code, and often does make a significant difference. (I understand it’s hard to support mutation, I’m just saying that this is something that would be very useful to have for general-purpose AD)

[Next page](https://discourse.julialang.org/t/ann-yota-jl-yet-another-reverse-mode-autodiff-package/21139.md?page=2)
