# Flux function fitting

**URL:** <https://discourse.julialang.org/t/flux-function-fitting/44475>\
**Category:** Machine Learning\
**Tags:** flux\
**Created:** [August 7, 2020, 4:45am UTC](https://discourse.julialang.org/t/flux-function-fitting/44475 "2020-08-07T04:45:37Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jstrube](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jstrube/32/525_2.png) [@jstrube](https://discourse.julialang.org/u/jstrube)\
**Post date:** [August 7, 2020, 4:45am UTC](https://discourse.julialang.org/t/flux-function-fitting/44475/1 "2020-08-07T04:45:37Z")

</div>

I’m trying to reproduce the results from this Tensorflow tutorial:

> **[Neural networks curve fitting | Lulu's blog](https://lucidar.me/en/neural-networks/curve-fitting-nonlinear-regression/)**
>
> Simple curve fitting with neural network / deep learning. Example, details and explanation of multi-layer neural-network nonlinear regression with TensorFlow.

The objective is a nonlinear regression. Sounds simple enough, but I’m struggling to reproduce the results from the web page. My attempt looks like this:

```julia
using Flux
using StatsPlots
using IterTools: ncycle
xvals = collect(Float32, range(-10.0, 10.0, length=1000))
xvals = reshape(xvals, (1, 1000))
yvals = 0.1f0.*xvals'.*cos.(xvals') + 0.1f0*randn(Float32, length(xvals))
scatter(xvals', yvals, ms=0.1, linewidth=0, markerstrokewidth=0, legend=nothing)
approx = Chain(
    Dense(1, 64),
    Dense(64, 64, relu),
    Dense(64, 64, relu),
    Dense(64, 1)
)
loss(x, y) = Flux.Losses.mse(x', y)
trainer = Flux.Data.DataLoader((xvals, yvals[:, 1]), shuffle=true)
Flux.train!(loss, params(approx), ncycle(trainer, 100), ADAM())
scatter(xvals', yvals, ms=0.1, linewidth=0, markerstrokewidth=0, legend=nothing)
plot!(xvals', approx(xvals)')

```

I’m using the same layers, activation functions and optimizer as the example (at least I think I am).  
Unfortunately, the result isn’t even close. In 100 epochs, the example fits the shape reasonably well, but the code I posted here basically just gives two straight lines, one for negative x-values and one for positive x-values.  
I could always try different layer structures, and different optimizers, but I feel like I’m missing something obvious, because this is a simple example that I should be able to reproduce. It works in Tensorflow, so there is no reason it shouldn’t work in Flux. Thank you for any suggestions on how to get closer to the Tensorflow performance.

Edit: I’ve also looked at the following posts, which contain useful information for how to solve the problem in general, but I couldn’t find an explanation for why the performance seems to be so different from that posted in the webpage.

> [@Regression in Flux.jl](https://discourse.julialang.org/t/regression-in-flux-jl/38733/3):
>
> Yep, that was it. Coming from another discipline, I thought “training” meant estimation. It’s very strange to me to have the training stop after one iteration…

> [@Generic Function to train NN w/ Flux](https://discourse.julialang.org/t/generic-function-to-train-nn-w-flux/37208/6):
>
> I suspect a lot of the issues here stem from the problem setup and not just Flux. You may want to start by examining the [condition number](https://en.wikipedia.org/wiki/Condition_number) of your data matrix. I found that k = 14992.546605226838 (the lower the condition number, the better conditioned your matrix). Such a large k indicates that for solving Ax = b with your data matrix, the solution will be very sensitive to errors. This makes iterative numerical methods prone to accumulating mistakes. A sign that this is happening is evident in t…

---

<div class="post-metadata">

**Author:** ![contradict](https://avatars.discourse-cdn.com/v4/letter/c/ac91a4/32.png) [@contradict](https://discourse.julialang.org/u/contradict)\
**Post date:** [August 7, 2020, 6:02am UTC](https://discourse.julialang.org/t/flux-function-fitting/44475/2 "2020-08-07T06:02:49Z")

</div>

One important typo, your loss function is evaluating x vs y with no intervening model, leading to the nice line you get. Try

```julia
loss(x, y) = Flux.Losses.mse(approx(x)', y)

```

And two less important differences.

1. Keras uses a batch size of 32 by [default](https://keras.rstudio.com/reference/fit.html), [`Flux.DataLoader`](https://fluxml.ai/Flux.jl/stable/data/dataloader/) does not batch unless you ask.

```julia
trainer = Flux.Data.DataLoader((xvals, yvals[:, 1]), shuffle=true, batchsize=32)

```

1. Keras uses `0.01` as the default initial [learning rate](https://keras.io/api/optimizers/adam/) for ADAM, Flux uses `0.001` default initial [learning rate](https://fluxml.ai/Flux.jl/stable/training/optimisers/#Flux.Optimise.ADAM)

```julia
Flux.train!(loss, params(approx), ncycle(trainer, 100), ADAM(0.01))

```

After these fixes, I see an acceptable fit.

---

<div class="post-metadata">

**Author:** ![jstrube](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jstrube/32/525_2.png) [@jstrube](https://discourse.julialang.org/u/jstrube)\
**Post date:** [August 7, 2020, 1:24pm UTC](https://discourse.julialang.org/t/flux-function-fitting/44475/3 "2020-08-07T13:24:16Z")

</div>

Urgh. I knew I was missing something obvious. Of course Flux can only train the network if the loss function actually depends on the network parameters…  
Thank you very much for taking a look and helping out!
