# Simple Flux model not learning

**URL:** <https://discourse.julialang.org/t/simple-flux-model-not-learning/30138>\
**Category:** Machine Learning\
**Tags:** flux\
**Created:** [October 21, 2019, 3:29pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138 "2019-10-21T15:29:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![rsteckel](https://avatars.discourse-cdn.com/v4/letter/r/45deac/32.png) [@rsteckel](https://discourse.julialang.org/u/rsteckel)\
**Post date:** [October 21, 2019, 3:29pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138/1 "2019-10-21T15:29:36Z")

</div>

I’ve been trying to get a feel for Flux by modeling a simple non-linear function (X^2).

At the end of each epoch, I print the model output for 3 (expecting 9), but it seems to converge to ~34

Can anyone see what I’m doing wrong?

```julia
using Flux
using Printf

model = Chain(
          Dense(1, 50),
          Dense(50, 1))

x = collect(-10:.1:10)'
y = x.^2
N = length(y)

loss(x, y) = Flux.mse(model(x), y)

opt = ADAM()
epochs = 15

ps = Flux.params(model)

@progress for epoch = 1:epochs  
  for i = 1:N
    gs = Flux.Tracker.gradient(() -> loss(x[:,i], y[i]), ps)
    Flux.Tracker.update!(opt, Tracker.Params(ps), gs) 
  end
  @printf "Epoch: %d 3^2 = %1.2f\n" epoch model([3]).data[1]
end

```

---

<div class="post-metadata">

**Author:** ![robsmith11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robsmith11/32/29641_2.png) [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Post date:** [October 21, 2019, 3:36pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138/2 "2019-10-21T15:36:25Z")

</div>

> [@rsteckel](#):
>
> x[:,i]

Isn’t `x` a 1-dimensional vector?

EDIT: Nevermind, missed the transpose.

---

<div class="post-metadata">

**Author:** ![robsmith11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robsmith11/32/29641_2.png) [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Post date:** [October 21, 2019, 4:14pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138/3 "2019-10-21T16:14:28Z")

</div>

I think the main thing you’re missing is a non-linear activation function. You can’t replicate the square of a number by simply taking a linear combination of itself.

By using `relu`, I got decent convergence after 100 epochs with your example.

I also tried using batch gradients, but convergence was very slow… about 10,000 epochs. There’s probably some tuning required to make it faster.

EDIT: Yep, `ADAM(0.1)` with batch gradients gets very good convergence by 1000 epochs and runs far faster than iterating through each data point.

---

<div class="post-metadata">

**Author:** ![PeterD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/peterd/32/10910_2.png) [@PeterD](https://discourse.julialang.org/u/PeterD)\
**Post date:** [October 21, 2019, 5:05pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138/4 "2019-10-21T17:05:59Z")

</div>

Like already stated you’ll need a non linear activation like a relu. Additionally I expect that the range of your input data is too large [-10,10] for training and too correlated between successive samples. I would suggest something along the lines of:

`x = randn(1, nr_of_samples_you_want)`

Hope this improves things a bit.

---

<div class="post-metadata">

**Author:** ![rsteckel](https://avatars.discourse-cdn.com/v4/letter/r/45deac/32.png) [@rsteckel](https://discourse.julialang.org/u/rsteckel)\
**Post date:** [October 21, 2019, 7:05pm UTC](https://discourse.julialang.org/t/simple-flux-model-not-learning/30138/5 "2019-10-21T19:05:21Z")

</div>

Ahh. I thought the default activation function was sigmoid. I see now that it’s an identity function.

(and using x = randn(1, nr\_of\_samples\_you\_want) helped also)

Thank you both.
