# Different behaviour between Flux.jl and Pytorch

**URL:** https://discourse.julialang.org/t/different-behaviour-between-flux-jl-and-pytorch/54941
**Category:** Machine Learning
**Tags:** machine-learning
**Created:** [February 9, 2021, 4:46pm UTC](https://discourse.julialang.org/t/different-behaviour-between-flux-jl-and-pytorch/54941 "2021-02-09T16:46:27Z")
**Posts on this page:** 1
**Showing post:** 2

<div class="post-metadata">

### Author: ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)
#### Post date: [February 9, 2021, 6:29pm UTC](https://discourse.julialang.org/t/different-behaviour-between-flux-jl-and-pytorch/54941/2 "2021-02-09T18:29:42Z")

</div>

> [@DeepQ](#):
>
> Another question: How can I trace the behaviour of gradient?

Yes, use `Flux.gradient` or `pullback` instead of `train!` so that you can analyze the gradients and model parameters before each update.

Another point to consider is that Flux and PyTorch initialize Dense layers differently by default. See [Initializing Flux weights the same as PyTorch? - #4 by DevJac](https://discourse.julialang.org/t/initializing-flux-weights-the-same-as-pytorch/54660/4). If you can verify that a) the initializations are similar, b) the outputs from each intermediate step of the forward pass are similar, and c) the gradients are similar, then I think the behaviour shouldn’t be much different between PyTorch and Flux. If there is a bug (which seems somewhat unlikely since you’re using a plain MLP on CPU), I would imagine it’s somewhere in the backwards pass (and thus will show up in the gradients).

---

_[View the full topic](https://discourse.julialang.org/t/different-behaviour-between-flux-jl-and-pytorch/54941)._
