# How do I make my custom (potentially stupid) model in Julia?

**URL:** <https://discourse.julialang.org/t/how-do-i-make-my-custom-potentially-stupid-model-in-julia/104464>\
**Category:** Machine Learning\
**Created:** [October 1, 2023, 1:31pm UTC](https://discourse.julialang.org/t/how-do-i-make-my-custom-potentially-stupid-model-in-julia/104464 "2023-10-01T13:31:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tarny\_GG\_Channie](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Tarny\_GG\_Channie](https://discourse.julialang.org/u/Tarny_GG_Channie)\
**Post date:** [October 1, 2023, 1:31pm UTC](https://discourse.julialang.org/t/how-do-i-make-my-custom-potentially-stupid-model-in-julia/104464/1 "2023-10-01T13:31:30Z")

</div>

Julia is said to be an expressive ML language where you do the obvious thing, putting in math expression and it works. But… how does it work exactly? To try it out, I want to implement my own (potentially stupid) idea for a language model.  
This idea is based on three basic assumptions.

1. Parallel computation between all nodes is good. (Like Transformer)
2. Low complexity is good (Like RNN to LSTM)
3. Attention is good (From GRU to transformer and so on).  
So, I proposed a (potentially stupid) idea based on binary attention.  
The main connections (inspired by wavenet, stride is 1, 2, 4, 8, …)  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/e/5e04dcee5cdde4eb17f0ba9a1b2fd0d038226567.png)  
Each connection is a binary attention.  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/2/a/2ae6b9805cc571636de7c1042b1f9c0162fa4468.png)

● If the left side is empty, set the sigmoid result to 1, meaning  
attending fully. If the right side is empty, set the sigmoid result to 0.  
● Go both directions, use this set of binary attention  
blocks to weight the result’s value. Now, it means that  
all tokens have attended both forward and backward!  
● Share weight at each layer. Do this until all tokens have inputs from all tokens from previous layers.  
● Repeat this process N times.

So, with this stupid idea, I want to try implementing it. How do I go about it?

---

<div class="post-metadata">

**Author:** ![tchebycheff](https://avatars.discourse-cdn.com/v4/letter/t/779978/32.png) [@tchebycheff](https://discourse.julialang.org/u/tchebycheff)\
**Post date:** [October 1, 2023, 2:29pm UTC](https://discourse.julialang.org/t/how-do-i-make-my-custom-potentially-stupid-model-in-julia/104464/2 "2023-10-01T14:29:43Z")

</div>

[Custom layers in Flux](https://fluxml.ai/Flux.jl/stable/models/advanced/)
