# What is the difference between Zygote vs ForwardDiff and ReverseDiff

**URL:** <https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789>\
**Category:** Machine Learning\
**Created:** [February 22, 2021, 5:41pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789 "2021-02-22T17:41:42Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![qft](https://avatars.discourse-cdn.com/v4/letter/q/b19c9b/32.png) [@qft](https://discourse.julialang.org/u/qft)\
**Post date:** [February 22, 2021, 5:41pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789/1 "2021-02-22T17:41:42Z")

</div>

As far as I can tell from the documentation, both ForwardDiff and ReverseDiff can already do automatic differentiation on arbitrary Julia functions. Also Zygote depends on ForwardDiff. What unique feature does Zygote brings to the picture?

---

<div class="post-metadata">

**Author:** ![anon74562486](https://avatars.discourse-cdn.com/v4/letter/a/c68b51/32.png) [@anon74562486](https://discourse.julialang.org/u/anon74562486)\
**Post date:** [February 23, 2021, 9:58pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789/2 "2021-02-23T21:58:18Z")

</div>

I am not an expert, so feel free to wait for answers from other users, they are much more experienced than I am.  
I am studying deep learning and with Zygote I can easly differentiate Flux neural network models.  
It is also much faster than the others you mentioned (I have tried it vs ReverseDiff with some simple feedforward networks).  
As far as I know (I’m a student) forward mode automatic differentiation is used when you have few parameters, otherwise you have to use reverse mode automatic differentiation.

* * *

> [@qft](#):
>
> both ForwardDiff and ReverseDiff can already do automatic differentiation on arbitrary Julia functions

No, all the libraries you mentioned have limitations.

* * *

> [@qft](#):
>
> Also Zygote depends on ForwardDiff.

Where did you read that?

* * *

For the study of deep learning I always use Zygote.  
Sorry for my english.

---

<div class="post-metadata">

**Author:** ![qft](https://avatars.discourse-cdn.com/v4/letter/q/b19c9b/32.png) [@qft](https://discourse.julialang.org/u/qft)\
**Post date:** [February 23, 2021, 10:43pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789/3 "2021-02-23T22:43:03Z")

</div>

Thank you for your kind response. You mentioned Zygote being faster than ReverseDiff. Is this generally the case? Are there circumstances that ForwardDiff/ReverseDiff performs faster than Zygote?

> [@anon74562486](#):
>
> No, all the libraries you mentioned have limitations.

The main reason I was making this post was that I would like know _what_ are the main limitations of ForwardDiff/ReverseDiff compared with Zygote (and vice versa). From the documentations ([Limitations of ForwardDiff · ForwardDiff](https://juliadiff.org/ForwardDiff.jl/latest/user/limitations.html) and [Limitations of ReverseDiff - ReverseDiff.jl](https://juliadiff.org/ReverseDiff.jl/limits/)) of these libraries, it seem the requirements are very lenient? One notable limitation for these two libraries is that they do not support mutation. But Zygote does not seem to support mutation either.

> [@anon74562486](#):
>
> Where did you read that?

ForwardDiff is listed as a dependency of Zygote in Project.toml ([https://github.com/FluxML/Zygote.jl/blob/6b89a068e40bad9673e163e9aee43f2bc4940242/Project.toml](https://github.com/FluxML/Zygote.jl/blob/6b89a068e40bad9673e163e9aee43f2bc4940242/Project.toml)).

---

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [February 23, 2021, 11:11pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789/4 "2021-02-23T23:11:50Z")

</div>

For a few-sentence summary of those (and several more) AD packages, see [https://juliadiff.org/](https://juliadiff.org/).

Zygote and ReverseDiff are both reverse-mode AD, but while ReverseDiff pushes custom types through your code to compute the backward pass (hence your code must be written to accept generic types), Zygote effectively rewrites the source code of your functions and works through more arbitrary code. For example, this fails:

```julia
ReverseDiff.gradient((x::Vector{Float64}) -> sum(x), ones(10))

```

but replacing ReverseDiff with Zygote works (of course in this trivial example its easy to make ReverseDiff work by just dropping that type annotation, but often its not this easy, especially if the code your differentiating is in someone else’s package).

The dependency of Zygote on ForwardDiff is just for a small piece used when broadcasting over CuArrays, Zygote is still reverse mode.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 23, 2021, 11:27pm UTC](https://discourse.julialang.org/t/what-is-the-difference-between-zygote-vs-forwarddiff-and-reversediff/55789/5 "2021-02-23T23:27:28Z")

</div>

> [@marius311](#):
>
> Zygote and ReverseDiff are both reverse-mode AD

To clarify, “reverse mode” AD is efficient when you have a functions f(x) with small number of outputs f\_i and many inputs x\_j (in computing \partial f\_i/\partial x\_j), i.e. for functions mapping x\in\mathbb{R}^m to f \in \mathbb{R}^n with n \ll m. (For example, in neural-network training where you want the derivative of one loss function (n=1) with respect to millions (m) of network parameters. (The “manual” application of such a technique is also known as an [adjoint method](https://math.mit.edu/~stevenj/18.336/adjoint.pdf), and in the neural-net case it is called backpropagation.)

In contrast, forward-mode AD (as in ForwardDiff.jl) is better when there is a small number of inputs and a _large_ number of outputs, i.e. when n \gg m, i.e. when you are computing _many_ functions of a _few_ variables. (It essentially corresponds to “manual” application of the chain rule in the most obvious way.)
