# How can I write a neural network using ForwardDiff.jl package?

**URL:** <https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408>\
**Category:** General Usage\
**Created:** [February 1, 2021, 7:00pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408 "2021-02-01T19:00:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![DeepQ](https://avatars.discourse-cdn.com/v4/letter/d/2bfe46/32.png) [@DeepQ](https://discourse.julialang.org/u/DeepQ)\
**Post date:** [February 1, 2021, 7:00pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/1 "2021-02-01T19:00:18Z")

</div>

I want to write a neural network using ForwardDiff.jl package but I can’t find any example.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [February 1, 2021, 7:06pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/2 "2021-02-01T19:06:58Z")

</div>

If number of parameters is much greater than number of outputs the you should use reverse mode AD.  
In a neural networks there normally are hundreds or thousands of parameters (the weights band biases), and 1 output (the loss).

With that said you can do this.  
Use `ForwardDiff.gradient`  
It will be faster than reverse mode if you have 5-10 parameters.

---

<div class="post-metadata">

**Author:** ![DeepQ](https://avatars.discourse-cdn.com/v4/letter/d/2bfe46/32.png) [@DeepQ](https://discourse.julialang.org/u/DeepQ)\
**Post date:** [February 1, 2021, 7:15pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/3 "2021-02-01T19:15:15Z")

</div>

I first tried to use Flux.jl to reproduce the results of a paper about normalizing flows. My Pytorch code works but the Flux code doesn’t work and it seems that Flux suffers from numerical problems. I don’t care about training time because I don’t think there would be a huge difference.  
It is strange that Flux.jl is the main DL package in Julia and it suffers from numerical problems!

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [February 1, 2021, 9:21pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/4 "2021-02-01T21:21:57Z")

</div>

> [@DeepQ](#):
>
> Flux suffers from numerical problems.

If you mean floating point truncation errors, then that is unlikely.  
AD doesn’t incur truncation errors.  
It may have bugs, if so you should open issues.  
but it really shouldn’t have round-off errors.  
Feel encourages to start another thread about that, someone might help you debug it.  
Or if you can get a vaguely minimal example of an error open an issue on GitHub.

> I don’t care about training time because I don’t think there would be a huge difference.

It will be orders of magnitude.  
To compute the gradient via forward mode involves basically running the code onces per parameter.  
Where as reverse mode is once per output (which is to say once).

Reverse mode does have a higher overhead, but unless the net is a tiny toy from the 90s, then it won’t dominate.

---

<div class="post-metadata">

**Author:** ![DeepQ](https://avatars.discourse-cdn.com/v4/letter/d/2bfe46/32.png) [@DeepQ](https://discourse.julialang.org/u/DeepQ)\
**Post date:** [February 2, 2021, 4:28am UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/5 "2021-02-02T04:28:33Z")

</div>

So it is better to use ReverseDiff.jl? My whole model has less than two thousand parameters and I don’t need to train it for long. The loss almost converges after 100 epochs.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [February 2, 2021, 5:13am UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/6 "2021-02-02T05:13:06Z")

</div>

I do not think that errors will be a problem. There is a bit outdated package implementing Masked Autoregressive flows here

> **[GitHub - gpapamak/maf: Masked Autoregressive Flow](https://github.com/gpapamak/maf)**
>
> Masked Autoregressive Flow. Contribute to gpapamak/maf development by creating an account on GitHub.

Also, we have written

> **[GitHub - pevnak/SumProductTransform.jl: An experimental implementation of...](https://github.com/pevnak/SumProductTransform.jl)**
>
> An experimental implementation of sum-product networks with dense unitary transformations in leaves - GitHub - pevnak/SumProductTransform.jl: An experimental implementation of sum-product networks ...

which uses “Dense” flows, where Dense matrix is optimized in its SVD form, which allows efficient calculation of Jacobian and inverse. We never had a problem with numerical stability.

Tomas

---

<div class="post-metadata">

**Author:** ![DeepQ](https://avatars.discourse-cdn.com/v4/letter/d/2bfe46/32.png) [@DeepQ](https://discourse.julialang.org/u/DeepQ)\
**Post date:** [February 2, 2021, 5:29am UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/7 "2021-02-02T05:29:42Z")

</div>

Try implementing Spline flows. I tried for months and it seems there is an unsolved bug in Flux.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [February 2, 2021, 9:47pm UTC](https://discourse.julialang.org/t/how-can-i-write-a-neural-network-using-forwarddiff-jl-package/54408/8 "2021-02-02T21:47:30Z")

</div>

> [@DeepQ](#):
>
> I tried for months and it seems there is an unsolved bug in Flux.

If bugs are not reported then they are not fixed.
