# PyTorch DataLoader equivalent for training large models with Flux

**URL:** https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763
**Category:** Machine Learning
**Tags:** flux
**Created:** [November 5, 2019, 6:25pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763 "2019-11-05T18:25:02Z")
**Posts on this page:** 17
**Page:** 1

<div class="post-metadata">

### Author: ![holylorenzo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/holylorenzo/32/34169_2.png) [@holylorenzo](https://discourse.julialang.org/u/holylorenzo)
#### Post date: [November 5, 2019, 6:25pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/1 "2019-11-05T18:25:03Z")

</div>

I’m trying to train a large computer vision model I built in Flux. Unfortunately the dataset does not fit into memory.

Does something similar to PyTorch’s `DataLoader` exist as a Julia package? Is there an ongoing effort on creating one? Has anyone else come across the need for something like it before?

I would be interested to hear your thoughts on this, thanks!

---

<div class="post-metadata">

### Author: ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)
#### Post date: [November 5, 2019, 9:24pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/2 "2019-11-05T21:24:47Z")

</div>

Not sure what data loader does but I am working on JDF.jl which allows each column to be loaded. I am developing methods allow chunk loading or random minibathc loading in v0.4. It’s no where near ready. But feel free to list down Ur requirements. I think just random minibatches?

---

<div class="post-metadata">

### Author: ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)
#### Post date: [November 5, 2019, 9:43pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/3 "2019-11-05T21:43:24Z")

</div>

I have found [MLDataUtils.jl](https://github.com/JuliaML/MLDataUtils.jl) pretty convenient for this kind of tasks.

---

<div class="post-metadata">

### Author: ![holylorenzo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/holylorenzo/32/34169_2.png) [@holylorenzo](https://discourse.julialang.org/u/holylorenzo)
#### Post date: [November 7, 2019, 12:08pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/4 "2019-11-07T12:08:59Z")

</div>

MLDataUtils.jl has a nice interface, similar to PyTorch’s `DataSet` and `DataLoader` with `nobs` and `getobs`.

To give some more background, PyTorch’s `DataLoader` basically wraps a data container and makes sure that you can get batches of observations quickly by distributing the load accross multiple threads. This is important especially for computer vision tasks as loading large images and performing expensive transformations are CPU-bound and can’t be precomputed because you there would only be enough memory for a few batches.

I suppose with Julia 1.3’s multi-threaded IO this might be a lot easier to implement? Has anyone done this, generally or especially in the context of machine learning?

---

<div class="post-metadata">

### Author: ![JohnnyChen94](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnnychen94/32/29979_2.png) [@JohnnyChen94](https://discourse.julialang.org/u/JohnnyChen94)
#### Post date: [November 9, 2019, 2:39pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/5 "2019-11-09T14:39:15Z")

</div>

> Unfortunately the dataset does not fit into memory.

`Dataloader` doesn’t do lazy-loading, it’s `Dataset` that does.

Although it’s not as convenient as `Dataset`, you can implement a function based on [MappedArrays.jl](https://github.com/JuliaArrays/MappedArrays.jl), for example:

```julia
using MappedArrays
using FileIO

function load_dataset(root)
    files = map(x->joinpath(root, x), readdir(root))
    return mappedarray(load, files)
end

root = "/Users/jc/Downloads/dataset"

# add ; in interactive environment to disable loading all files
dataset = load_dataset(root);

```

Image is not read from disk until it’s used, and that’s the trick `Dataset` does to save your memory. However, the same image is read from disk multiple times, which would be slower than directly reading from memory. It’s a tradeoff between computational time and memory space.

---

<div class="post-metadata">

### Author: ![PeterD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/peterd/32/10910_2.png) [@PeterD](https://discourse.julialang.org/u/PeterD)
#### Post date: [November 10, 2019, 3:37pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/6 "2019-11-10T15:37:20Z")

</div>

I’ve been working on a framework on top of Knet called Photon and as a proof of concept have implemented Dataset/Dataloader functionality (including treading). So perhaps that can serve as inspiration?

You can find the code in subdirectory src/data in

[https://github.com/neurallayer/Photon.jl](https://github.com/neurallayer/Photon.jl)

---

<div class="post-metadata">

### Author: ![holylorenzo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/holylorenzo/32/34169_2.png) [@holylorenzo](https://discourse.julialang.org/u/holylorenzo)
#### Post date: [November 12, 2019, 9:46am UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/7 "2019-11-12T09:46:56Z")

</div>

This looks very nice!  
Does it run on Julia 1.3 (rc4) already? If I’m not mistaken threaded IO did not work reliably before until incl. 1.2.

---

<div class="post-metadata">

### Author: ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)
#### Post date: [November 12, 2019, 10:15am UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/8 "2019-11-12T10:15:29Z")

</div>

Looks real nice. Love how the chain takes care of input size for me so I don’t have to specify that in the chain

---

<div class="post-metadata">

### Author: ![PeterD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/peterd/32/10910_2.png) [@PeterD](https://discourse.julialang.org/u/PeterD)
#### Post date: [November 19, 2019, 10:13pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/9 "2019-11-19T22:13:57Z")

</div>

In general threading does work well, even with IO involved. However some external packages are not yet thread-safe. One bug I found for example is that ImageMagick.jl works well, but not when used through FileIO.

---

<div class="post-metadata">

### Author: ![PeterD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/peterd/32/10910_2.png) [@PeterD](https://discourse.julialang.org/u/PeterD)
#### Post date: [November 19, 2019, 10:15pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/10 "2019-11-19T22:15:44Z")

</div>

Thanks! After using MXNet for a long time, really got addicted to this feature. Before that I always was to lazy to calculate the output sizes (like in PyTorch) and just used the debugger to figure it out 😉

---

<div class="post-metadata">

### Author: ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)
#### Post date: [November 19, 2019, 10:22pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/11 "2019-11-19T22:22:46Z")

</div>

> [@PeterD](#):
>
> just used the debugger to figure it

Same. Framework is meant to make it easy right?

---

<div class="post-metadata">

### Author: ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)
#### Post date: [November 19, 2019, 10:24pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/12 "2019-11-19T22:24:10Z")

</div>

In your opinion, what are some reasons why someone would choose Julia ecosystem over others?

---

<div class="post-metadata">

### Author: ![PeterD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/peterd/32/10910_2.png) [@PeterD](https://discourse.julialang.org/u/PeterD)
#### Post date: [November 20, 2019, 2:16pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/13 "2019-11-20T14:16:27Z")

</div>

My assessment is similar to that of Google in this article ([https://github.com/tensorflow/swift/blob/master/docs/WhySwiftForTensorFlow.md](https://github.com/tensorflow/swift/blob/master/docs/WhySwiftForTensorFlow.md)). The difference being that Google selected the lesser option out of the two final languages (IMHO) when they decided to go with Swift.

1. With \> 1 million mobile developers, I cannot see their benefit to morph Swift into a ML/numerical language. Often design decisions will be conflicting. So Swift for TensorFlow could stay a fork for a long time or even forever. Julia has both the right features and community.

2. Fully static type checking is a burden at the beginning of data science projects (exploration phase). I think Julia strikes a better balance here, although a bit more compile time checking would be welcome.

3. Many of the better tools, as to be expected, for Swift are MacOs based (btw a platform with limited NVidia support). Julia has for such a new language excellent tooling in place already for all major platforms. Using Juno on a daily basis and replaces the two IDE syndrome (Notebooks and PyCharm).

But I have proven wrong more often than I like to admit, so who knows 😉

BTW I also looked at Kotlin Native quickly. They did some very cool bindings with TensorFlow and PyTorch as a proof of concept and I was very impressed with the results. But in the end also not as suited as Julia for datascience and numerical computing IMHO.

---

<div class="post-metadata">

### Author: ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)
#### Post date: [November 20, 2019, 5:55pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/14 "2019-11-20T17:55:42Z")

</div>

I have been using extensively this one `https://github.com/pevnak/DataIterators.jl`, but it uses processes rather than threads (as at the time of writing threads were not available). I would like to consolidate this effort, as I would like to write this kind of thing for threads as well.

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [November 24, 2019, 2:47pm UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/15 "2019-11-24T14:47:48Z")

</div>

I also have an embryo of something similar  
[https://github.com/baggepinnen/DiskDataProviders.jl](https://github.com/baggepinnen/DiskDataProviders.jl)  
It works well for what I’m doing, but could certainly be made more general.

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [November 27, 2019, 9:30am UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/16 "2019-11-27T09:30:18Z")

</div>

See also [[ANN] LengthChannels - Buffered iterators for machine learning](https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566)

---

<div class="post-metadata">

### Author: ![terasakisatoshi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/terasakisatoshi/32/9187_2.png) [@terasakisatoshi](https://discourse.julialang.org/u/terasakisatoshi)
#### Post date: [November 8, 2020, 8:58am UTC](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/17 "2020-11-08T08:58:03Z")

</div>

How about [holylorenzo](https://discourse.julialang.org/u/holylorenzo)’s repository?

> [@\[ANN\] DataLoaders.jl (alpha) - basically PyTorch's parallel \`DataLoader\`](https://discourse.julialang.org/t/ann-dataloaders-jl-alpha-basically-pytorchs-parallel-dataloader/35963):
>
> Happy to give a first look at a (as of yet unregistered) package [DataLoaders.jl](https://github.com/lorenzoh/DataLoaders.jl) that has a similar API to PyTorch’s DataLoader. See examples in the [README.md](https://github.com/lorenzoh/DataLoaders.jl/blob/master/README.md) I wrote this primarily to support deep learning pipelines that load images, apply heavy preprocessing, and batch the samples. To do this without slowing down the training this has to be done on multiple threads and non-blockingly. A few months ago [I asked](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/14) if something like this already exists in the Julia ecosystem, but did not find anyt…
