# Dense Layers in Attention Don't Train

**URL:** <https://discourse.julialang.org/t/dense-layers-in-attention-dont-train/117190>\
**Category:** Machine Learning\
**Created:** [July 18, 2024, 2:57pm UTC](https://discourse.julialang.org/t/dense-layers-in-attention-dont-train/117190 "2024-07-18T14:57:16Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Helectus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/helectus/32/47475_2.png) [@Helectus](https://discourse.julialang.org/u/Helectus)\
**Post date:** [July 18, 2024, 2:57pm UTC](https://discourse.julialang.org/t/dense-layers-in-attention-dont-train/117190/1 "2024-07-18T14:57:17Z")

</div>

I am trying to create a transformer model using Flux.jl from scratch. I have an encoder struct that has an attention struct, which includes Q, K, V, O dense layers. My problem is that Dense.weights and Dense.bias don’t seem to change after training. They seem to be the same after a lot of epochs.

I have added the  
Flux.@functor AttentionLayer  
as well as Flux.@functor Encoder etc.

The overall model seems to be getting trained, with the MSE loss decreasing, but while other parameters seem to change, the Dense layers in the attention don’t seem to be affected.

Anyone who could help me?

---

<div class="post-metadata">

**Author:** ![dmolina](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmolina/32/5246_2.png) [@dmolina](https://discourse.julialang.org/u/dmolina)\
**Post date:** [July 20, 2024, 4:19pm UTC](https://discourse.julialang.org/t/dense-layers-in-attention-dont-train/117190/2 "2024-07-20T16:19:15Z")

</div>

Welcome.

First, could you give a Minimum example to help you?

Second, maybe you should interested in [GitHub - chengchingwen/Transformers.jl: Julia Implementation of Transformer models](https://github.com/chengchingwen/Transformers.jl), It has an interesting documentation.
