# Do I understand Enzyme properly?

**URL:** <https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760>\
**Category:** Specific Domains\
**Tags:** adjoint, autodiff, enzyme\
**Created:** [April 21, 2023, 9:10pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760 "2023-04-21T21:10:01Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 21, 2023, 9:10pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/1 "2023-04-21T21:10:01Z")

</div>

I think I finally have a grasp on how Enzyme works, but I would like an expert to validate this.

Consider a possibly mutating vector function f(x, y) which outputs a scalar value z. We denote by x\_a and y\_a the contents of x and y _after execution_. Enzyme works with a Jacobian including all of these variables (defined by blocks):

$J = \begin{pmatrix}  
\partial x\_a / \partial x & \partial x\_a / \partial y \  
\partial y\_a / \partial x & \partial y\_a / \partial y \  
\partial z / \partial x & \partial z / \partial y \  
\end{pmatrix}$

In forward mode, if we give `Duplicated` arguments:

1. we provide a tuple of values (x, y) and their tangents v = (\dot{x}, \dot{y}) as shadows
2. `autodiff` computes Jv = (\dot{x\_a}, \dot{y\_a}, \dot{z})
3. it puts \dot{x\_a} and \dot{y\_a} in the shadows we provided
4. it returns \dot{z}

In reverse mode, if we give `Duplicated` arguments:

1. we provide a tuple of values (x, y) and their adjoints v = (\bar{x\_a}, \bar{y\_a}) as shadows
2. … which is completed with a default output adjoint \bar{z} = 1 (I guess?)
3. `autodiff` computes v^\top J = (\bar{x}, \bar{y})
4. it puts \bar{x} and \bar{y} in the shadows we provided (EDIT: no, it’s an addition)
5. it returns `nothing`

Do I have that right? I think such a description with the full Jacobian would have helped me in the documentation (coming from ChainRules)

Sources:

- [Basics · Enzyme.jl](https://enzyme.mit.edu/julia/stable/generated/autodiff/)
- [https://enzyme.mit.edu/julia/stable/pullbacks/](https://enzyme.mit.edu/julia/stable/pullbacks/)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 21, 2023, 9:33pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/2 "2023-04-21T21:33:13Z")

</div>

I think you have the Jacobian products in step (2/3) backwards. Forwards-mode AD is equivalent to Jacobian–vector products Jv (it evaluates the chain rule from _right-to-left_, or equivalently from _inputs-to-outputs_). Backwards-mode AD is equivalent to vector–Jacobian products v^TJ (it evaluates the chain rule from _left-to-right_, or equivalently from _outputs-to-inputs_).

This [video on principles of AD](https://www.youtube.com/watch?v=UqymrMG-Qi4) by @mohamed82008 may be helpful.

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 21, 2023, 9:35pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/3 "2023-04-21T21:35:14Z")

</div>

Thanks for spotting the switch, I corrected it!

Do you have any clue as to the default seeding of the output adjoint in reverse mode Enzyme? It’s the part I’m really unsure about cause in ChainRules we give it explicitly to the VJP

That’s also why I asked specifically about Enzyme, because in ChainRules we don’t care about storage being updated unless it influences the output (so the x\_a and y\_a don’t really exist there)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 21, 2023, 9:40pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/4 "2023-04-21T21:40:42Z")

</div>

> [@gdalle](#):
>
> Do you have any clue as to the default seeding of the output adjoint in Enzyme? It’s the part I’m really unsure about cause in ChainRules we give it explicitly to the VJP

Just write down the chain rule for computing f(x) = L(F(G(x)):

f'(x) = L'(F(G(x)) F'(G(x)) G'(x)

where each of the terms L', F', G' is a Jacobian matrix and x, F, G are vectors. Imagine that L is a loss function with a scalar output. Then the Jacobian L' is a row vector v^T where v = \nabla L. Reverse-mode multiplies (\nabla f)^T = L' F' G' = (v^T F') G' left-to-right, so that you are always doing vector–Jacobian products (never matrix–matrix products). (This is precisely the case where reverse-mode is best: many inputs x and a single scalar output L.)

We have a short course _Matrix Calculus_ at MIT that covers this and more: [GitHub - mitmath/matrixcalc: MIT IAP short course: Matrix Calculus for Machine Learning and Beyond · GitHub](https://github.com/mitmath/matrixcalc)

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 21, 2023, 9:44pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/5 "2023-04-21T21:44:41Z")

</div>

I’m sorry, my question wasn’t clear. I’m used to computing VJPs and JVPs where you declare explicitly the whole vector v. My impression is that in Enzyme, that’s not the case:

- in forward mode, the user does specify \dot{x} and \dot{y}
- but in reverse mode they only specify \bar{x\_a} and \bar{y\_a}, which means the remaining component \bar{z} must come from somewhere: presumably it’s 1 (that’s what makes mathematical sense and also what my tests seem to suggest) but I’d love a confirmation cause I didn’t find anything about it in the docs

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 21, 2023, 9:50pm UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/6 "2023-04-21T21:50:00Z")

</div>

> [@stevengj](#):
>
> We have a short course _Matrix Calculus_ at MIT that covers this and more: [GitHub - mitmath/matrixcalc: MIT IAP short course: Matrix Calculus for Machine Learning and Beyond](https://github.com/mitmath/matrixcalc)

By the way, I don’t know if Alan told you but I wrote a new homework last year on autodiff with ChainRules for his Julia class, it may be useful for this one too!  
[https://gdalle.github.io/JuliaComputationSolutions/hw2\_solutions.html](https://gdalle.github.io/JuliaComputationSolutions/hw2_solutions.html)

---

<div class="post-metadata">

**Author:** ![wsmoses](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wsmoses/32/26497_2.png) [@wsmoses](https://discourse.julialang.org/u/wsmoses)\
**Post date:** [April 22, 2023, 4:29am UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/7 "2023-04-22T04:29:29Z")

</div>

@gdalle your understanding looks mostly correct.

My minor clarification is that for reverse mode it doesn’t put \bar{x} into the storage you provide it +='s it into the storage you provide (which we refer to as shadow memory and/or the dval in the duplicated).

In combined reverse mode (the default) the default output adjoint is indeed one. This is related to why right now reverse-mode does not permit duplicated returns. The reason being that the augmented forward pass in reverse mode will return a shadow data structure for the return, which we will need to += intowith the output adjoint seed. For arbitrary data structures the question of how to do so is more fuzzy (since what if the output of the forward pass doesn’t match the user provided adjoint structure, since it will be runtime generated). For ease we presently just support Const or Active returns as a result.

However, if you want to seed with arbitrary values (including Duplicated returns like arrays, trees, etc) you have options, a couple of which I’ll put here.

1. Return by reference. Therefore you have a duplicated whre you would initialize the duplicated with the adjoint you want to propagate. Enzyme will propagate that value (and also zero it when it takes it out – this is necessary to ensure the proper functioning of loops since if you store to the same memory location twice, only the last stored value is the one you want to propagate to).
2. Use split reverse mode. First run the augmented forward pass, which returns a shadow data structure. += your adjoint into the shadow (which can be an array, matrix, linked list, etc).

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [April 22, 2023, 8:43am UTC](https://discourse.julialang.org/t/do-i-understand-enzyme-properly/97760/8 "2023-04-22T08:43:45Z")

</div>

Thanks @wsmoses, I’ll take time to digest your answer (which is the one I was looking for) and come back if I have more questions 😈

Special shoutout to @stevengj, who seems to answer every single question on this forum without ever losing patience 🙏
