# How to freeze a single weight of a layer?

**URL:** <https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279>\
**Category:** General Usage\
**Tags:** flux\
**Created:** [March 31, 2021, 7:37am UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279 "2021-03-31T07:37:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![March-08](https://avatars.discourse-cdn.com/v4/letter/m/96bed5/32.png) [@March-08](https://discourse.julialang.org/u/March-08)\
**Post date:** [March 31, 2021, 7:37am UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/1 "2021-03-31T07:37:55Z")

</div>

Is there a way to freeze one single weight instead of an entire layer ?  
`Flux.delete!` function is able to freeze just the layer and not one single parameter.

```julia
 function micronet()
    simplenet = gpu(Chain(
      Dense(2, 2 relu),
      Dense(2,1),
    ));

```

```julia
function dummy_data_micronet(train::Int64, test::Int64)
        xtrain = rand(train,2) |> cpu
        ytrain = rand(train) |> cpu
        xtest = rand(test,2) |>cpu
        ytest = rand(test) |> cpu
        data = [(xtrain', ytrain')] |> cpu

        return data,xtest',ytest'
    end

```

```julia
model = SimpleNet.micronet()
data,xtest,ytest = DataCostumizer.dummy_data_micronet(1000,100)
loss(x, y) = Flux.mse(model(x),y)
ps = params(model)
opt = ADAM()

```

```julia
evalcb = () -> println("removed edges : ",MagnitudePruning.compute_zero_entries(model), "\n loss : ",loss(xtest,ytest))
@Flux.epochs 10 Flux.train!(loss, ps, data, opt,cb = throttle(evalcb, 10))    
delete!(ps,model[1].W[2])
@Flux.epochs 10 Flux.train!(loss, ps, data, opt,cb = throttle(evalcb, 10)) 

```

The same for `Flux.trainable`

---

<div class="post-metadata">

**Author:** ![March-08](https://avatars.discourse-cdn.com/v4/letter/m/96bed5/32.png) [@March-08](https://discourse.julialang.org/u/March-08)\
**Post date:** [April 2, 2021, 10:23am UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/3 "2021-04-02T10:23:44Z")

</div>

Thank you for your answer!  
I have tried your method, but it seems that in this way none of the parameters updates.

```julia
model = Chain(
      Flux.Dense(2, 2 ,relu),
      Flux.Dense(2,1),
      softmax
    )

# %%
data,xtest,ytest = DataCostumizer.dummy_data_micronet(10000,1000)

loss(x, y) = Flux.mse(model(x),y)
p = params(model)
opt = ADAM()

epochs = 1

indices_to_update = [i for i in 1:length(p[1]) if i ≠ 4] # indices of p[1] you want to update
p |> println

for epoch in 1:epochs
    for (x, y) in data
        gr = gradient(p) do 
            loss(x, y)
        end 
        for i in 1:length(p)
            if i ≠ 1
                Flux.update!(opt, p[i], gr[p[i]])
            else
                Flux.update!(opt, 
                             p[i][indices_to_update], 
                             gr[p[i]][indices_to_update])
            end
        end
    end
end

p |> println  

```

The output are :

```julia
Params([Float32[-0.8721409 1.1937331; -0.71039486 1.121799], Float32[0.0, 0.0], Float32[-0.14334783 0.29841754], Float32[0.0]])

```

and again :

```julia
Params([Float32[-0.8721409 1.1937331; -0.71039486 1.121799], Float32[0.0, 0.0], Float32[-0.14334783 0.29841754], Float32[0.0]])

```

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [April 2, 2021, 4:51pm UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/4 "2021-04-02T16:51:46Z")

</div>

The code above still removes an entire parameter. The easiest way to freeze a single connection/weight is to zero or mask out the corresponding index in the gradient array:

```julia
...

for epoch in 1:epochs
    for (x, y) in data
        gr = gradient(p) do 
            loss(x, y)
        end

        # remove offending gradient
        gr[model[1].W][2] = 0
        Flux.update!(opt, p, gr)
    end
end

```

---

<div class="post-metadata">

**Author:** ![anon74562486](https://avatars.discourse-cdn.com/v4/letter/a/c68b51/32.png) [@anon74562486](https://discourse.julialang.org/u/anon74562486)\
**Post date:** [April 2, 2021, 6:29pm UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/5 "2021-04-02T18:29:36Z")

</div>

My answer is wrong, ToucheSir is right.  
I used that method to update parameters other than 0 and didn’t realize it was incorrect there as well.  
I was stupid, I apologize.

---

<div class="post-metadata">

**Author:** ![March-08](https://avatars.discourse-cdn.com/v4/letter/m/96bed5/32.png) [@March-08](https://discourse.julialang.org/u/March-08)\
**Post date:** [April 7, 2021, 9:21am UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/6 "2021-04-07T09:21:22Z")

</div>

Thanks for your answer!  
But I was wondering, this method works only if the updates depend solely on the gradient right ?  
For example with the ADAM optimizer each single weight changes even if the gradient is set to zero because it doesn’t depend only on the value of the gradient.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [April 7, 2021, 4:24pm UTC](https://discourse.julialang.org/t/how-to-freeze-a-single-weight-of-a-layer/58279/7 "2021-04-07T16:24:11Z")

</div>

> [@March-08](#):
>
> For example with the ADAM optimizer each single weight changes even if the gradient is set to zero because it doesn’t depend only on the value of the gradient.

Not so. If you look at the [definition of ADAM in Flux](https://github.com/FluxML/Flux.jl/blob/master/src/optimise/optimisers.jl#L175-L182), you’ll notice that the parameter update will be 0 if the gradient, `mt` and `vt` are all 0 as well. Since we know the gradient for the frozen weight is 0 and the optimizer params are initialized with `zeros`, we can guarantee that weight will never be updated because all operations are applied element-wise.
