# Flux seq2seq

**URL:** <https://discourse.julialang.org/t/flux-seq2seq/13783>\
**Category:** Machine Learning\
**Tags:** question\
**Created:** [August 20, 2018, 7:36pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783 "2018-08-20T19:36:43Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [August 20, 2018, 7:36pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/1 "2018-08-20T19:36:43Z")

</div>

I don’t know whether this is the right place to ask but I’m trying to code a seq2seq model in Flux and I’ve got a couple of questions.

The encoder creates the hidden state for the decoder, how can I pass this state to the `GRU` in the decoder? Is this possible without a custom-built layers?

Secondly, how can I use this model on the gpu? It seems like `|> gpu` doesn’t work on line 9 because of the indexing?  
Also `result` is of undefined size, so this can’t be ran on my gpu?

Thanks,  
Jules

```julia
model = function(seq, voc_size, max_length)
    seq = onehotbatch(seq, dictionary_fr[:, 1])
    seq = emb_layer_fr*seq
    
    #split seq in it's columns:
    seq = [seq[:, i] for i in 1:size(seq)[2]]
    
    #encoder
    x = GRU(300, 256).(seq)[end]
    x = Dense(256, 300)(x)
    
    #decoder
    result = Vector{Any}(undef, 0)
    input = onehot(1, 1:voc_size) #<BOS>
    for i in 1:max_length
        input = emb_layer_nl*input
        output = Chain(GRU(300, 300), Dense(300, voc_size))(input)
        append!(result, output)
        input = onehot(argmax(output), 1:voc_size)
        input.ix == 3 ? break : continue
    end
    return(result)
end

```

---

<div class="post-metadata">

**Author:** ![dellison](https://avatars.discourse-cdn.com/v4/letter/d/ce73a5/32.png) [@dellison](https://discourse.julialang.org/u/dellison)\
**Post date:** [August 22, 2018, 10:37pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/2 "2018-08-22T22:37:33Z")

</div>

I suggest that you take a look at [model-zoo/1-model.jl at master · FluxML/model-zoo · GitHub](https://github.com/FluxML/model-zoo/blob/master/text/phonemes/1-model.jl) for an example of how an encoder/decoder model can be implemented in Flux. With the code you’re showing here, it looks like you’re creating a new copy of every single layer every time you call this function (and even throwing away and recreating the decoder layers each trip through the loop), and I’m guessing that this probably is not what you have in mind for your model’s behavior.

Regarding the GPU question, I haven’t actually used this functionality myself, but the documentation for it is [here](http://fluxml.ai/Flux.jl/stable/gpu.html). My understanding is that you’ll need to make sure to call Flux’s `gpu` function on both your model’s weights and its inputs to make sure that everything has the appropriate GPU types.

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [August 28, 2018, 3:43pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/3 "2018-08-28T15:43:07Z")

</div>

Thanks,  
I’ve changed my code quite a bit, but now, my model doesn’t return correct translations.

When I overfit my model with just one sentence pair, it gives the correct output, but when I try to overfit it with 2 (or more) sentence pairs, it returns a mix of the two sentences. It looks like the model just picks the most prevalent word from one of the sentences.

[This](https://nextjournal.com/a/KEbhYTEtFefLrZks2SFrR) is the link to a nextjournal notebook with my code and data (I can’t get it running there though).

I would highly appreciate if anyone could take a quick look and tell me what I’m doing wrong or link a julia implementation of machine translation. Chances are I’m making an easy mistake since I’m a complete beginner.

Thanks,  
Jules

---

<div class="post-metadata">

**Author:** ![jekbradbury](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jekbradbury/32/953_2.png) [@jekbradbury](https://discourse.julialang.org/u/jekbradbury)\
**Post date:** [August 28, 2018, 10:04pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/4 "2018-08-28T22:04:14Z")

</div>

I’m not sure there’s a comprehensive Julia implementation of seq2seq for MT anywhere, although I’ve talked about working on one and I may be able to look at your code soon if I get some time.

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [September 1, 2018, 2:09pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/5 "2018-09-01T14:09:49Z")

</div>

I finally found the problem, my input didn’t suit the encoder network. I’m going to upload my code as soon as possible.

---

<div class="post-metadata">

**Author:** ![Nakul\_Tiruviluamala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nakul_tiruviluamala/32/11228_2.png) [@Nakul\_Tiruviluamala](https://discourse.julialang.org/u/Nakul_Tiruviluamala)\
**Post date:** [October 22, 2018, 2:47pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/6 "2018-10-22T14:47:07Z")

</div>

@merckxiaan, I spoke to the nextjournal team to figure out why I couldn’t access your notebook (I am interested in seq2seq in Julia). It turns out that you did not publish the journal so as a result no one else can see it. I’d be greatly obliged if you would!

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [October 22, 2018, 3:14pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/7 "2018-10-22T15:14:43Z")

</div>

Hello @Nakul_Tiruviluamala,  
After I posted my last message, I abandoned this project since I’m a beginner and I was making mistake after mistake.  
Since then however, I’m trying to implement Pytorch’s tutorial on seq2seq machine translation, I’m trying to follow the tutorial as close as possible. Now I believe I’m stuck due to a bug in Flux which prohibits me from concatenating a transposed array ([https://github.com/FluxML/Flux.jl/issues/378](https://github.com/FluxML/Flux.jl/issues/378)). I probably also made a lot of mistakes here and there. It would be great if you could have a look at the code and let me know your thougts/questions.  
[Nextjournal](https://nextjournal.com/a/KS7QNE3vRMYhQMSwDKiFX/edit)

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [October 22, 2018, 3:18pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/8 "2018-10-22T15:18:12Z")

</div>

> [Nextjournal](https://nextjournal.com/a/KS7QNE3vRMYhQMSwDKiFX/edit)

You need to click **publish** , to get a shareable link! That’s just your internal edit link 😉

 ![publish](https://global.discourse-cdn.com/julialang/original/3X/9/0/902c3cb36894e8ba10bd6202499545c10a10e0bd.gif)

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [October 22, 2018, 3:19pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/9 "2018-10-22T15:19:22Z")

</div>

whoops!

> **[Pytorch seq2seq machine translation in Flux](https://nextjournal.com/jules/pytorch-seq2seq-machine-translation-in-flux)**

---

<div class="post-metadata">

**Author:** ![Nakul\_Tiruviluamala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nakul_tiruviluamala/32/11228_2.png) [@Nakul\_Tiruviluamala](https://discourse.julialang.org/u/Nakul_Tiruviluamala)\
**Post date:** [October 23, 2018, 4:00pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/10 "2018-10-23T16:00:26Z")

</div>

Hi @merckxiaan, I’ll definitely be looking at it. I am a beginner as well!

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [October 26, 2018, 8:34pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/11 "2018-10-26T20:34:30Z")

</div>

Nice @Nakul_Tiruviluamala ,  
I’ve just uploaded my code to a github gist because I’m running in an error with Julia. [Issue](https://github.com/JuliaLang/julia/issues/29818)  
Even though my code crashes, the loss does decline for a few steps…  
Perhaps you could try to reproduce this error?

Jules

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [October 31, 2018, 8:21pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/12 "2018-10-31T20:21:32Z")

</div>

Nevermind, I’ve started over once more. Does anyone spot something wrong with my encoder, decoder or attention layer, for some reason, when I train my model the loss get’s stuck and the model predicts some frequent words.

```julia
struct Encoder
    embedding
    rnn
end
Encoder(voc_size::Int, h_size::Int) = Encoder(
    param(Flux.glorot_uniform(h_size, voc_size)),
    GRU(h_size, h_size))
function (e::Encoder)(x; dropout=0)
    x = e.embedding*x
    x = Dropout(dropout)(x)
    x = e.rnn(x)
    return(x)
end
Flux.@treelike Encoder

struct Decoder
    embedding
    attention
    rnn
    output
end
Decoder(h_size, voc_size) = Decoder(
    param(Flux.glorot_uniform(h_size, voc_size)),
    Attention(h_size),
    GRU(h_size*2, h_size),
    Dense(h_size, voc_size, relu))
function (d::Decoder)(x, encoder_outputs; dropout=0)
    x = d.embedding * x
    x = Dropout(dropout)(x)
    decoder_state = d.rnn.state
    context = d.attention(encoder_outputs, decoder_state)
    x = d.rnn([x; context])
    x = softmax(d.output(x))
end
Flux.@treelike Decoder

struct Attention
    linear
end
Attention(h_size::Int) = Attention(Dense(2*h_size, 1, tanh))
function (a::Attention)(encoder_outputs, decoder_state)
    weights = []
    results = []
    for word in encoder_outputs
        weight = a.linear([word; decoder_state])
        push!(weights, weight)
    end
    weights = softmax(vcat(weights...))
    return sum([encoder_outputs[i].*weights[i, :]' for i in 1:size(weights, 1)])
end
Flux.@treelike Attention

```

---

<div class="post-metadata">

**Author:** ![merckxiaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/merckxiaan/32/4697_2.png) [@merckxiaan](https://discourse.julialang.org/u/merckxiaan)\
**Post date:** [January 5, 2019, 12:52pm UTC](https://discourse.julialang.org/t/flux-seq2seq/13783/13 "2019-01-05T12:52:55Z")

</div>

I’ve made some progress and put all my code, with some explanations, in a [notebook](https://github.com/merckxiaan/flux-seq2seq/blob/master/seq2seq%20in%20flux.ipynb). The model does seem to learn something… more often than not, the subject of the sentence is correct, but the remaining words are gibberish.

Also I notice a big difference in performance with different hyperparameters but I’m not sure how I could choose the optimal ones.

I’d really appreciate someone providing me with some feedback.

Thanks,  
Jules
