# Params not getting updated during training

**URL:** <https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985>\
**Category:** New to Julia\
**Tags:** flux\
**Created:** [October 8, 2020, 2:13pm UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985 "2020-10-08T14:13:02Z")\
**Posts on this page:** 6\
**Page:** 2

<div class="post-metadata">

**Author:** ![m\_scorpion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m_scorpion/32/18509_2.png) [@m\_scorpion](https://discourse.julialang.org/u/m_scorpion)\
**Post date:** [October 9, 2020, 6:24pm UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/21 "2020-10-09T18:24:37Z")

</div>

i started from a minimal code again and it is stuck training after first round  
i guess i have to go back to mx and tf for now  
i found it , when dataloader batchsize is more than 1 , it gets like this

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [October 9, 2020, 8:50pm UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/22 "2020-10-09T20:50:11Z")

</div>

This seems like a misunderstanding of Flux vs TF/MXnet APIs. Could you please post the actual minimal code example you tried (with dummy data generation if need be) and also the Python code you’re trying to translate? I ask because the issue seems to most likely be in the data preprocessing or loading side rather than the model definition or training.

---

<div class="post-metadata">

**Author:** ![m\_scorpion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m_scorpion/32/18509_2.png) [@m\_scorpion](https://discourse.julialang.org/u/m_scorpion)\
**Post date:** [October 10, 2020, 12:08am UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/23 "2020-10-10T00:08:51Z")

</div>

i have posted the simple code above  
im still working on different ways to do this

---

<div class="post-metadata">

**Author:** ![m\_scorpion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m_scorpion/32/18509_2.png) [@m\_scorpion](https://discourse.julialang.org/u/m_scorpion)\
**Post date:** [October 11, 2020, 11:01am UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/24 "2020-10-11T11:01:18Z")

</div>

here is the minimal code :

```julia
using Flux
model = Chain(Dense(10,1))
x=rand([1,2,3,4,5,6,7,8,9,0],(10,10000))
y=rand([1,2,3,4,5,6,7,8,9,0],(1,10000))
function losss(x,y)
    return Flux.mae(model(x),y)
end
optimiser = Flux.Descent(0.01)
train_loader = Flux.Data.DataLoader((x,y),batchsize=1)
Flux.@epochs 10 Flux.train!(losss,params(model),train_loader,optimiser,
    cb = Flux.throttle(() -> println(losss(x,y)),10))

```

the problem raises after first epoch , all the weights and biases in model becomes NaN and as the result model’s output becomes NaN ,also the loss result ,hence the training fails to continue…

---

<div class="post-metadata">

**Author:** ![theogf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/theogf/32/1987_2.png) [@theogf](https://discourse.julialang.org/u/theogf)\
**Post date:** [October 11, 2020, 3:08pm UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/25 "2020-10-11T15:08:09Z")

</div>

This could be due to an overflow, `Descent()` does not protect you from exploding gradients and such. You could either clip your gradients `Flux.Optimise.Optimiser(ClipGradient(10.0), Descent(0.01))` or use a momentum based optimiser like `ADAM()`.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [October 11, 2020, 9:12pm UTC](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985/26 "2020-10-11T21:12:49Z")

</div>

I would also be surprised if this _didn’t_ result in NaNs if you wrote something similar in PyTorch, mostly because there’s no activation function like sigmoid/tanh/etc. to reduce the size of the outputs. If you try reducing the magnitude of your inputs (either by dividing everything by, say 100 or using a `rand` variant that samples from `[0, 1)`), the network should be less susceptible to spitting out NaNs. You may also want to try a larger batch size: it will “smooth out” the gradient updates and thus also reduce the possibility of overflowing into NaNs.

[Previous page](https://discourse.julialang.org/t/params-not-getting-updated-during-training/47985.md?page=1)
