# Gradient in simple regression not working

**URL:** <https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839>\
**Category:** Machine Learning\
**Tags:** flux\
**Created:** [August 22, 2019, 8:55am UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839 "2019-08-22T08:55:20Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ali\_Vahdati](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ali_vahdati/32/38715_2.png) [@Ali\_Vahdati](https://discourse.julialang.org/u/Ali_Vahdati)\
**Post date:** [August 22, 2019, 8:55am UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839/1 "2019-08-22T08:55:20Z")

</div>

I tried to replicate a simple regression model from Flux’s documentation with a slightly larger dataset. But the gradient descent seems to always diverge. A manual implementation of the algorithm works well.

I would appreciate if anyone can point the problem.

Here is my code using Flux:

```julia

using Flux
using Flux.Tracker
using Flux.Tracker: update!
using RDatasets

trees = dataset("datasets", "trees");

X = Matrix(trees[!, [:Girth,:Height]])
y = trees[!, :Volume]
n = length(y)
nfeatures = size(X, 2)

W = rand(nfeatures)
b = rand()

predict(X, W, b) = X*W .+ b

function loss(X, y, W, b)
  ŷ = predict(X, W, b)
  sum((y .- ŷ).^2)
end

loss(X, y, W, b)
# ~ 27139.27213905282

W = param(W)
b = param(b)
gs = Tracker.gradient(() -> loss(X, y, W, b), params(W, b))
# Update the parameter and reset the gradient
update!(W, -0.0003gs[W])
update!(b, -0.0003gs[b])
loss(X, y, W, b)
# ~ 2.6542565871928763e8 (tracked)

```

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [August 22, 2019, 9:27am UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839/2 "2019-08-22T09:27:18Z")

</div>

Are you sure your stepsize is good and have the correct sign?

---

<div class="post-metadata">

**Author:** ![Ali\_Vahdati](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ali_vahdati/32/38715_2.png) [@Ali\_Vahdati](https://discourse.julialang.org/u/Ali_Vahdati)\
**Post date:** [August 22, 2019, 9:30am UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839/3 "2019-08-22T09:30:52Z")

</div>

It has the correct sign, following the example in the docs. This is the size I use in a manual implementation of the model and it works. Nevertheless, I have checked an order of magnitude smaller step size and it still diverges.

---

<div class="post-metadata">

**Author:** ![MikeInnes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikeinnes/32/3656_2.png) [@MikeInnes](https://discourse.julialang.org/u/MikeInnes)\
**Post date:** [September 2, 2019, 2:52pm UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839/4 "2019-09-02T14:52:44Z")

</div>

The best way forward is probably to try and gradually make the flux implementation look more like the manual version, or vice versa, until you figure out what the difference is between the two. You’re already calculating gradients directly, so perhaps check that lines up with your other version?

---

<div class="post-metadata">

**Author:** ![Ali\_Vahdati](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ali_vahdati/32/38715_2.png) [@Ali\_Vahdati](https://discourse.julialang.org/u/Ali_Vahdati)\
**Post date:** [September 3, 2019, 10:40am UTC](https://discourse.julialang.org/t/gradient-in-simple-regression-not-working/27839/5 "2019-09-03T10:40:02Z")

</div>

Thanks for your reply.  
I changed the loss function to the following:

```julia
function loss(X, y, W, b)
  ŷ = predict(X, W, b)
  sum((y .- ŷ).^2) / (2*length(y))
end

```

And the gradients are correct now. However, I am not sure why taking the average of sums of squares is important for the gradients to work.
