# \[ANN\] LengthChannels - Buffered iterators for machine learning

**URL:** <https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566>\
**Category:** Package Announcements\
**Tags:** package, announcement\
**Created:** [November 27, 2019, 9:28am UTC](https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566 "2019-11-27T09:28:17Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [November 27, 2019, 9:28am UTC](https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566/1 "2019-11-27T09:28:17Z")

</div>

There has recently been many threads talking about buffered data loaders for machine learning workflows, e.g. [https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/](https://discourse.julialang.org/t/pytorch-dataloader-equivalent-for-training-large-models-with-flux/).  
I have been writing a few different solutions for this throughout the last 6 months, and have found that a rather low level is what I prefer. I made the package [LengthChannels.jl](https://github.com/baggepinnen/LengthChannels.jl) which is a modest extension of `Base.Channel` that endows the channel with a length for iteration. This allows the channel to be used effectively as a dataset by Flux etc.

The [README](https://github.com/baggepinnen/LengthChannels.jl) contains an example of a buffered batch iterator for Flux CNNs.

The package requires julia v1.3 and allows the population of the channel to take place on a **separate thread**. This way, data loading and potential pre-processing can be done by the CPU while the GPU is training your model. The data can not be put on the GPU by the channel as one is not allowed to mess with the GPU by a separate thread.

I also made the package [DiskDataProviders.jl](https://github.com/baggepinnen/DiskDataProviders.jl) which is a bit more complicated and may be useful when reading the data from disk takes longer time than a single training step. It was created with Flux in mind and is quite a lot less flexible than a `LengthChannel`.

---

<div class="post-metadata">

**Author:** ![zgornel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zgornel/32/217487_2.png) [@zgornel](https://discourse.julialang.org/u/zgornel)\
**Post date:** [November 28, 2019, 4:13pm UTC](https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566/2 "2019-11-28T16:13:25Z")

</div>

This looks awesome, cheers.

---

<div class="post-metadata">

**Author:** ![DoktorMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/doktormike/32/2736_2.png) [@DoktorMike](https://discourse.julialang.org/u/DoktorMike)\
**Post date:** [November 28, 2019, 4:42pm UTC](https://discourse.julialang.org/t/ann-lengthchannels-buffered-iterators-for-machine-learning/31566/3 "2019-11-28T16:42:11Z")

</div>

Nice work!
