# Parallel data loading to GPU arrays

**URL:** <https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321>\
**Category:** Machine Learning\
**Tags:** gpu, parallel, gpuarrays, data, flux\
**Created:** [June 1, 2018, 3:49am UTC](https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321 "2018-06-01T03:49:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [June 1, 2018, 3:49am UTC](https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321/1 "2018-06-01T03:49:28Z")

</div>

Anyone have pointers to loading data into CuArrays with background tasks while the main thread or process is busy training a Flux model?

Flux seems to have the same if not more flexibility than pytorch, so they fall in the same niche, and I would like to try out Flux instead for my next project.

However my experience is that for the kinds of loosely structured data that such flexibility helps the most with, loading and packaging the data for the model to consume quickly becomes the bottleneck. pytorch provides multiprocess [DataLoaders](https://pytorch.org/docs/master/data.html#torch.utils.data.DataLoader) for this. How would I do the same in the Julia ecosystem?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [June 1, 2018, 4:11am UTC](https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321/2 "2018-06-01T04:11:46Z")

</div>

Does `@async` not do it?

---

<div class="post-metadata">

**Author:** ![stratospark](https://avatars.discourse-cdn.com/v4/letter/s/c68b51/32.png) [@stratospark](https://discourse.julialang.org/u/stratospark)\
**Post date:** [January 30, 2019, 1:52am UTC](https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321/3 "2019-01-30T01:52:18Z")

</div>

Wondering if there are any good beginner tutorials about this?

I’m coming from the Tensorflow/Keras world, where we can easily handle larger than memory datasets with Dataset iterators.

Super excited about the possibilities of Julia/Flux, but having trouble making the jump from MNIST-sized toy datasets to large ones.

---

<div class="post-metadata">

**Author:** ![kolia](https://avatars.discourse-cdn.com/v4/letter/k/c68b51/32.png) [@kolia](https://discourse.julialang.org/u/kolia)\
**Post date:** [January 30, 2019, 3:07am UTC](https://discourse.julialang.org/t/parallel-data-loading-to-gpu-arrays/11321/4 "2019-01-30T03:07:54Z")

</div>

There is [this QueuePool](https://github.com/FluxML/Metalhead.jl/blob/sf/training/training/ImageNet/queuepool.jl) referenced from [Flux PR #450](https://github.com/FluxML/Flux.jl/pull/450#issuecomment-435025441).

Haven’t tried it, don’t know if it’s in a working state, and don’t have pointers to example usage, maybe @staticfloat does?
