# Comparation with Flux leads to odd results in Flux

**URL:** https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839
**Category:** Machine Learning
**Tags:** question, flux
**Created:** [July 10, 2020, 11:25am UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839 "2020-07-10T11:25:16Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [July 10, 2020, 11:25am UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/1 "2020-07-10T11:25:16Z")

</div>

Hello,  
I’m developing a toy ML library ([BetaML](https://github.com/sylvaticus/BetaML.jl)) to learn a bit about ML algorithms (I’m pretty a newbie).

I tested BetaML with a bike sharing demand forecast example and compared it with Flux:  
[https://github.com/sylvaticus/BetaML.jl/blob/master/notebooks/NN%20-%20Bike%20sharing%20demand%20forecast%20(daily%20db).ipynb](https://github.com/sylvaticus/BetaML.jl/blob/master/notebooks/NN%20-%20Bike%20sharing%20demand%20forecast%20(daily%20db).ipynb)

Runnable Binder notebook: [Binder](https://mybinder.org/v2/gh/sylvaticus/BetaML.jl/master?filepath=notebooks%25NN%20-%20Bike%20sharing%20demand%20forecast%20(daily%20db).ipynb)

However, when I use the same model structure, data, training algorithms and hyperparameters, I experience strange behaviour in Flux, e.g. the data predicted by Flux for the training sample seems to be truncated. Also, predictions seems not to move so much as with the BetaML results:

**BetaML output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/d/5d6230218ad24857772a9393cbeea0f4597c7db7.png)  
**Flux output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/c/d/cd7c2b41470f6f78bc43cfe42e2ec2b74e985fff.png)  
**BetaML output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/0/5/055f5d271bb26e5fbab8ef126c3d36718b89a500.png)  
**Flux output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/4/2/42d0b54b3b4294ce0f184069d20cd81cb089928b.png)  
**BetaML output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/1/a/1ac1323d643b306061abe735b7ea86a62de36eae.png)  
**Flux output:**  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/8/b/8bd5a8fa43f49f44d4d3b7a8da96f16f3b45b531.png)

Note that also BetaML results tend to underestimate the high demand levels observed in the validation period, but in a less pronounced way that Flux.

I am wondering what cause this difference… weight initialisation ?

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [July 10, 2020, 1:12pm UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/2 "2020-07-10T13:12:29Z")

</div>

Are you using relu activation functions? It looks like the output has been clamped in a way typical for relu. If this is the case, the problem tends to go away with further training or using some strategy to make it easier to find a better minimum, such as residual connections etc.

You can also try to normalize your data prior to train g so that it has mean zero and variance 1,it also helps with this problem.

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [July 10, 2020, 1:29pm UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/3 "2020-07-10T13:29:19Z")

</div>

Thank you, I am using sigmoid for the hidden layer and identity for the output one:

```julia

# Defining the net model and load it with data...
Flux_nn = Chain(Dense(23,12,Flux.sigmoid),
                Dense(12,1,identity))
loss(x, y) = Flux.mse(Flux_nn(x), y)
ps = Flux.params(Flux_nn)
nndata = Flux.Data.DataLoader(xtrain', ytrain', batchsize=8)

```

What I found strange is that using instead an other library - but with the same parameters, including batch size, optimizer and number of epochs - I don’t get this effect.

I will try increasing the epochs. But it could also be that Flux doesn’t do by default a certain number of tricks, like Xavier weight initialisation, or random sampling of the batches… I did notice that Flux has a philosophy not to provide default arguments/optimisations. I understand it, but for newcomers could be useful to have for example an optimizer and a loss function by default.

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [July 10, 2020, 1:31pm UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/4 "2020-07-10T13:31:08Z")

</div>

No Flux does not do any random sampling for you, it’s all up to the iterator you pass it. The weight initialization should be one of the standard ones though.

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [July 10, 2020, 1:34pm UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/5 "2020-07-10T13:34:22Z")

</div>

I would still normalize the data. Since the data you are predicting is in the 1000s, the initial gradients will drive all the activation functions to saturation before the linear output layer has caught up. By normalizing, you’ll probably have a much faster convergence and to a better minimum.

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [July 10, 2020, 1:39pm UTC](https://discourse.julialang.org/t/comparation-with-flux-leads-to-odd-results-in-flux/42839/6 "2020-07-10T13:39:23Z")

</div>

Yes, actually both X and Y are scaled [in the script](https://github.com/sylvaticus/BetaML.jl/blob/master/notebooks/NN%20-%20Bike%20sharing%20demand%20forecast%20(daily%20db).ipynb) (for Y they are just divided by 1000, as if I normalise to mean 0, s.d. 1 I may have some negative demand when I rescale them back).

Edit: you was right, to get similar results it was enought to load the data with shuffling: `nndata = Flux.Data.DataLoader(xtrain', ytrain', batchsize=8,shuffle=true) ` (in BetaML I do it by default unless you opt-off with `sequential=true`)
