# Constrain weights and biases to be positive

**URL:** <https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381>\
**Category:** Machine Learning\
**Tags:** question, flux\
**Created:** [January 23, 2023, 7:37am UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381 "2023-01-23T07:37:15Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![marco\_menarini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_menarini/32/46117_2.png) [@marco\_menarini](https://discourse.julialang.org/u/marco_menarini)\
**Post date:** [January 23, 2023, 7:37am UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/1 "2023-01-23T07:37:15Z")

</div>

Is it possible to use a custom constrained optimizer or to modify the gradients to only allow for positive weights and biases?

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [January 24, 2023, 8:51pm UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/2 "2023-01-24T20:51:17Z")

</div>

I want to say this has been discussed before on Discourse, but not being able to find a thread I’d suggest looking at [Feature request: Modifying Dense Layer to accommodate kernel/bias constraints and kernel/bias regularisation · Issue #1389 · FluxML/Flux.jl · GitHub](https://github.com/FluxML/Flux.jl/issues/1389). Note that we are moving away from this “implicit” `Params` model into storing both model weights and gradients in proper structures. If you’d like something more futureproof, have a look at [Home · Optimisers.jl](https://fluxml.ai/Optimisers.jl/dev/#An-optimisation-rule).

---

<div class="post-metadata">

**Author:** ![marco\_menarini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_menarini/32/46117_2.png) [@marco\_menarini](https://discourse.julialang.org/u/marco_menarini)\
**Post date:** [January 24, 2023, 9:31pm UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/3 "2023-01-24T21:31:42Z")

</div>

Thanks. I tried the regularization path but I couldn’t make it work, I am looking into the custom optimizer now 🙂

---

<div class="post-metadata">

**Author:** ![Rasmus\_Hoier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rasmus_hoier/32/24036_2.png) [@Rasmus\_Hoier](https://discourse.julialang.org/u/Rasmus_Hoier)\
**Post date:** [January 27, 2023, 1:04pm UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/4 "2023-01-27T13:04:54Z")

</div>

An alternative approach is to introduce an activation function for the weights. That way the weights “as applied” are non-negative, but they posses a real valued learnable “latent state”.

```julia
using Flux
       
m = Dense(3, 2, relu)
x = rand(Float32, 3, 5)
(a::Dense)(x::AbstractVecOrMat, g) = a.σ.(g.(a.weight)*x .+ g.(a.bias))

y1 = m(x) # normal forward pass
y2 = m(x, relu) # forward pass with non-negative weights

```

---

<div class="post-metadata">

**Author:** ![marco\_menarini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marco_menarini/32/46117_2.png) [@marco\_menarini](https://discourse.julialang.org/u/marco_menarini)\
**Post date:** [January 27, 2023, 6:29pm UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/5 "2023-01-27T18:29:01Z")

</div>

Oooh that is a good idea, now maybe stupid question…

I don’t see any problem using a normal gradient descent, but do you think it would create some problemsusing a method requiring an Hessian (stochastic LBFGS) for training since it depends on the information on the previous gradient? (that was the main reason why I was looking if it was possible to have a constrained optimizer 🙂

---

<div class="post-metadata">

**Author:** ![Rasmus\_Hoier](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rasmus_hoier/32/24036_2.png) [@Rasmus\_Hoier](https://discourse.julialang.org/u/Rasmus_Hoier)\
**Post date:** [January 30, 2023, 8:38am UTC](https://discourse.julialang.org/t/constrain-weights-and-biases-to-be-positive/93381/6 "2023-01-30T08:38:02Z")

</div>

I don’t think it should be an issue, but I might be missing something.

But in general nonnegative weights make a network a lot less expressive. You are limiting the weights to a single orthant. I think both approaches have drawbacks

- Optimizer approach: How to efficiently project to the feasible domain.

- weight activation approach: How to avoid dying or saturated synapses (ReLU and Sigmoid case respectively).
