# Uploading vector of vectors to GPU in flux.jl

**URL:** <https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552>\
**Category:** Machine Learning\
**Tags:** question\
**Created:** [September 21, 2021, 9:27pm UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552 "2021-09-21T21:27:48Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![mys\_721tx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mys_721tx/32/28238_2.png) [@mys\_721tx](https://discourse.julialang.org/u/mys_721tx)\
**Post date:** [September 21, 2021, 9:27pm UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552/1 "2021-09-21T21:27:48Z")

</div>

Hello all,

I am training a many-to-many RNN. My input and output data are both `Vector{Vector{Vector{Float32}}}`. The jld file of the input and output is about 3GB. I checked the model-zoo and did not find anything example on uploading RNN sequences to the GPU. I tried to upload the whole thing by using `data_train = zip(us_train, ys_train) |> gpu`. This ends up using all my GPU memory. Should I convert the input to a tensor of some sort? What is the best way to upload those sequence data to the GPU?

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [September 22, 2021, 4:13pm UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552/2 "2021-09-22T16:13:07Z")

</div>

The best way to do so is in minibatches, just like you would for any other data or framework. Unlike other layers, however, this will be a vector of arrays to represent the sequence dimension. In other words, each minibatch will look like `[(features x batch size) x sequence length]`. [Recurrence · Flux](https://fluxml.ai/Flux.jl/stable/models/recurrence/#Working-with-sequences) talks all about what kinds of data the built-in RNNs can work with.

---

<div class="post-metadata">

**Author:** ![mys\_721tx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mys_721tx/32/28238_2.png) [@mys\_721tx](https://discourse.julialang.org/u/mys_721tx)\
**Post date:** [September 23, 2021, 4:39am UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552/3 "2021-09-23T04:39:02Z")

</div>

I have played around with `DataLoader` and it takes a `[features x sequence length x batch size]` tensor.

The following snippet can run but the loss is way off. I will keep playing around and see what’s wrong.

```nohighlight
train_loader = Flux.DataLoader((data=us_train, label=ys_train), batchsize=5)

for (x, y) in train_loader
    x = x |> gpu
    y = y |> gpu

    Flux.train!(loss, ps, zip(x, y), opt, cb=evalcb)
end

```

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [September 23, 2021, 2:48pm UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552/4 "2021-09-23T14:48:35Z")

</div>

`x` and `y` are already batched, so zipping them feeds through samples 1 at a time. `Flux.train!` is not really flexible enough for this, so I’d recommend using a [custom training loop](https://fluxml.ai/Flux.jl/stable/training/training/#Custom-Training-loops) so that you’re able to transform the data into an RNN-compatible from before feeding it to the model.

---

<div class="post-metadata">

**Author:** ![mys\_721tx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mys_721tx/32/28238_2.png) [@mys\_721tx](https://discourse.julialang.org/u/mys_721tx)\
**Post date:** [October 1, 2021, 5:55pm UTC](https://discourse.julialang.org/t/uploading-vector-of-vectors-to-gpu-in-flux-jl/68552/5 "2021-10-01T17:55:52Z")

</div>

This is the training function I ended up having.

```nohighlight
function seq_batch_train!(loss, ps, data, opt; cb = () -> ())
    local training_loss
    cb = Flux.Optimise.runall(cb)

    x, y = data
    x = x |> gpu
    y = y |> gpu

    gs = Flux.gradient(ps) do
        training_loss = loss(x, y)
        return training_loss
    end
    @show(training_loss)
    Flux.Optimise.update!(opt, ps, gs)
    cb()
end

```

And here is the training loop with epoches.

```nohighlight
for epoch in 1:1
    @show(epoch)
    for batch in train_loader
        seq_batch_train!(loss, ps, batch, opt, cb = evalcb)
    end
    @save "model_$(now())_epoch-$epoch.bson" m opt
end

```

The trick is in the loss function:

```nohighlight
function loss(x, y)
    Flux.reset!(m)
    y_pred = [(m(x[:, xi, :]) - y[:, xi, :]).^2 for xi in axes(x, 2)] |> sum |> sum
    y_pred / length(x)
end

```

`hcat()` is very slow on GPU and therefore the MSE is manually calculated instead of using `mse()` from Flux.

Currently batch size of 10 uses about 8GB of GPU memory. I can see some tweaks to get it work with longer time series.
