# Approximating a Quadratic Function with Flux

**URL:** <https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034>\
**Category:** Machine Learning\
**Created:** [May 9, 2019, 2:57pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034 "2019-05-09T14:57:09Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![mboutros](https://avatars.discourse-cdn.com/v4/letter/m/ecc23a/32.png) [@mboutros](https://discourse.julialang.org/u/mboutros)\
**Post date:** [May 9, 2019, 2:57pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/1 "2019-05-09T14:57:09Z")

</div>

I’m trying to get familiar with neural networks and Flux by estimating a series of simple models. First, I can successfully estimate a linear model using Flux:

```julia
using Plots
using Flux
using Flux: @epochs

gridsize = 100;
dgp(x) = -12x+3;
X = collect(range(0,stop=10,length=gridsize));
Y = dgp.(X);

data = []
for i in 1:length(X)
    push!(data, ([X[i]], Y[i]))
end

model = Chain(Dense(1,1))
loss(x, y) = Flux.mse(model(x), y)
opt = Descent(0.01)
ps = Flux.params(model)
@epochs 10 Flux.train!(loss, ps, data, opt)

# Plot.
plot(X,[Y model(X').data'],label=["DGP" "Model"])

```

With only a few iterations, the model does a pretty good job:

![21%20AM](https://global.discourse-cdn.com/julialang/original/3X/d/9/d96af3864999ca72c8384cfc2eb07b52837f3c34.png)

I’m running into problems trying to approximate a quadratic function. The code is largely the same:

```julia
using Plots
using Flux
using Flux: @epochs

gridsize = 100;
dgp(x) = x^2;
X = collect(range(0,stop=10,length=gridsize));
Y = dgp.(X);

data = []
for i in 1:length(X)
    push!(data, ([X[i]], Y[i]))
end

Q = 10;
model = Chain(Dense(1,Q,σ),
    Dense(Q,1,identity));

loss(x, y) = Flux.mse(model(x), y)
opt = Descent(0.01)
para = Flux.params(model)
@epochs 10 Flux.train!(loss, para, data, opt)

# Plot.
plot(X,[Y model(X').data'],label=["DGP" "Model"])

```

Theoretically, I should be able to represent the function `f(x) = x^2` over my compact grid, and 10 hidden layers (i.e., `Q = 10` in my code) should be sufficient for a fairly good approximation. Running this code, however, generates a very “flat” model:

![07%20AM](https://global.discourse-cdn.com/julialang/original/3X/8/8/8827ba9d5bda149ba62d64206ceb4907f05e6fbc.png)

I’ve tried different activation functions, changing the speed of the gradient descent, and a few other things, but I’m wondering if I’m doing something wrong within Flux. Thanks in advance for any help!

---

<div class="post-metadata">

**Author:** ![evanfields](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/evanfields/32/1744_2.png) [@evanfields](https://discourse.julialang.org/u/evanfields)\
**Post date:** [May 9, 2019, 3:30pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/2 "2019-05-09T15:30:18Z")

</div>

If I’m reading the Flux docs correctly, the `Dense` function’s second argument is just the number of outputs:

```julia
Flux.Dense — Type.
Dense(in::Integer, out::Integer, σ = identity)
Creates a traditional Dense layer with parameters W and b.

y = σ.(W * x .+ b)

```

So your

```julia
model = Chain(Dense(1,Q,σ),
    Dense(Q,1,identity));

```

doesn’t have `Q` hidden layers, it’s just two layers?

---

<div class="post-metadata">

**Author:** ![mboutros](https://avatars.discourse-cdn.com/v4/letter/m/ecc23a/32.png) [@mboutros](https://discourse.julialang.org/u/mboutros)\
**Post date:** [May 9, 2019, 3:33pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/3 "2019-05-09T15:33:51Z")

</div>

Sorry, that’s a typo. I actually have one hidden layer with 10 nodes; the second “layer” just linearly combines the nodes. This should should be enough for a good approximation to a simple quadratic function. I’d also tried increasing the number of hidden layers, but that doesn’t help with the “flatness,” either.

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [May 9, 2019, 5:18pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/4 "2019-05-09T17:18:49Z")

</div>

The following small changes lead to a pretty good fit:

```julia
Q = 10;
model = Chain(Dense(1,Q,tanh),
    Dense(Q,1,identity));

loss(x, y) = Flux.mse(model(x), y)
opt = ADAM(.001)
para = Flux.params(model)
@epochs 500 Flux.train!(loss, para, data, opt)

```

---

<div class="post-metadata">

**Author:** ![charlesll](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/charlesll/32/3255_2.png) [@charlesll](https://discourse.julialang.org/u/charlesll)\
**Post date:** [May 9, 2019, 9:15pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/5 "2019-05-09T21:15:06Z")

</div>

So this is not related to Flux but to the activation function.

You also can try putting more neurons, 10 is small even for a simple quadratic unction… Those days it is better to build large/deep networks than shallow ones.

I will soon put an example [here](https://github.com/charlesll/legendary-machine-learning) of some tests I did, showing than a 3 layers - 100 neurons per layer network performs better at predicting a sigmoid function than a simple 1 layer - 20 neurons network, despite being trained on 40 points only…

---

<div class="post-metadata">

**Author:** ![mboutros](https://avatars.discourse-cdn.com/v4/letter/m/ecc23a/32.png) [@mboutros](https://discourse.julialang.org/u/mboutros)\
**Post date:** [May 22, 2019, 5:51pm UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/6 "2019-05-22T17:51:13Z")

</div>

Thank you all for the helpful replies. This answer helped a lot! In fact, if I leave everything in my original code the same, but decrease the learning rate, it also works very well.

---

<div class="post-metadata">

**Author:** ![charlesll](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/charlesll/32/3255_2.png) [@charlesll](https://discourse.julialang.org/u/charlesll)\
**Post date:** [May 30, 2019, 8:24am UTC](https://discourse.julialang.org/t/approximating-a-quadratic-function-with-flux/24034/7 "2019-05-30T08:24:57Z")

</div>

In relation to my message, [here](https://github.com/charlesll/legendary-machine-learning/blob/master/Test_MLsize_Flux_Julia.ipynb) is finally a first version of the notebook!! We see that a moderately deep network performs better than the shallow one…
