# \[ANN\] DataLoaders.jl (alpha) - basically PyTorch's parallel \`DataLoader\`

**URL:** <https://discourse.julialang.org/t/ann-dataloaders-jl-alpha-basically-pytorchs-parallel-dataloader/35963>\
**Category:** Package Announcements\
**Tags:** flux, machine-learning\
**Created:** [March 14, 2020, 11:13am UTC](https://discourse.julialang.org/t/ann-dataloaders-jl-alpha-basically-pytorchs-parallel-dataloader/35963 "2020-03-14T11:13:17Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![holylorenzo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/holylorenzo/32/34169_2.png) [@holylorenzo](https://discourse.julialang.org/u/holylorenzo)\
**Post date:** [March 14, 2020, 11:13am UTC](https://discourse.julialang.org/t/ann-dataloaders-jl-alpha-basically-pytorchs-parallel-dataloader/35963/1 "2020-03-14T11:13:17Z")

</div>

Happy to give a first look at a (as of yet unregistered) package [`DataLoaders.jl`](https://github.com/lorenzoh/DataLoaders.jl) that has a similar API to PyTorch’s `DataLoader`.

See examples in the [`README.md`](https://github.com/lorenzoh/DataLoaders.jl/blob/master/README.md)

I wrote this primarily to support deep learning pipelines that load images, apply heavy preprocessing, and batch the samples. To do this without slowing down the training this has to be done on multiple threads and non-blockingly.

A few months ago [I asked](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/30763/14) if something like this already exists in the Julia ecosystem, but did not find anything that suited my needs, so I decided to roll my own.

Note: in a recent release of `Flux.jl` [a `DataLoader` was added](https://github.com/FluxML/Flux.jl/pull/1051), but as far as I can see, the implementation is neither parallel nor non-blocking.

I’d be happy about feedback, criticism, and whether you find the package helpful.
