# Convention/package for chunking array stream?

**URL:** <https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395>\
**Category:** General Usage\
**Tags:** array\
**Created:** [January 8, 2019, 7:19pm UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395 "2019-01-08T19:19:12Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zach\_Christensen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zach_christensen/32/7220_2.png) [@Zach\_Christensen](https://discourse.julialang.org/u/Zach_Christensen)\
**Post date:** [January 8, 2019, 7:19pm UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395/1 "2019-01-08T19:19:12Z")

</div>

Is there a convention or package for reading/writing in an array by chunks? I understand I could do it myself piecewise using `read`, but it seems that this could quickly become complicated when taking into account multiple file types, element types being read (e.g., colors, etc). Maybe something like DataStreams but generalized to multidimensional arrays.

---

<div class="post-metadata">

**Author:** ![oatlzzvztd](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oatlzzvztd/32/3285_2.png) [@oatlzzvztd](https://discourse.julialang.org/u/oatlzzvztd)\
**Post date:** [January 8, 2019, 7:33pm UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395/2 "2019-01-08T19:33:54Z")

</div>

I’ve found [Blobs.jl](https://github.com/RelationalAI-oss/Blobs.jl) useful in the past.

---

<div class="post-metadata">

**Author:** ![ssfrr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ssfrr/32/3736_2.png) [@ssfrr](https://discourse.julialang.org/u/ssfrr)\
**Post date:** [January 8, 2019, 7:46pm UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395/3 "2019-01-08T19:46:27Z")

</div>

[SampledSignals.jl](https://github.com/JuliaAudio/SampledSignals.jl) defines semantics for reading/writing multichannel chunks of audio samples. The basic idea is that you have `SampleSource` and `SampleSink` abstract types that you can `read` and `write` from, respectively. These chunks are represented as `SampleBuf`s (though if you don’t want to store samplerate then `Array{T}` would work just as well). When you open an audio file or device, you get a concrete subtype that knows how to do the encoding/decoding.

@samoconnor has done some really nice work reviewing the conventions for `read`/`write` APIs in Base:  
[https://github.com/JuliaLang/julia/issues/24526#issuecomment-431567472](https://github.com/JuliaLang/julia/issues/24526#issuecomment-431567472)

Then there’s also the question of whether you really want to expose your streams with a `read`/`write` API like these or whether it makes more sense to use `iterate` over the elements, or a more FRP-style API using something like [`Observables.jl`](https://github.com/JuliaGizmos/Observables.jl) or [Signals.jl](https://github.com/TsurHerman/Signals.jl) (there are maybe others, too).

Historically in the audio world software works in chunks for efficiency, particularly when signal graphs get wired together dynamically with function pointers, so you don’t want to pay the indirection overhead on every sample. With Julia it may not be necessary to operate chunk-wise, because the per-sample operations can be coalesced and optimized jointly.

---

<div class="post-metadata">

**Author:** ![Zach\_Christensen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zach_christensen/32/7220_2.png) [@Zach\_Christensen](https://discourse.julialang.org/u/Zach_Christensen)\
**Post date:** [January 9, 2019, 4:32am UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395/4 "2019-01-09T04:32:41Z")

</div>

I have a pretty limited knowledge of the breadth of IO streams so this is all very helpful.

> [@oatlzzvztd](#):
>
> I’ve found [Blobs.jl](https://github.com/RelationalAI-oss/Blobs.jl) useful in the past.

This is really interesting, as I’ve seen several packages trying to do something similar. This seems to be the easiest approach for streaming simple structs I’ve seen so far.

> [@ssfrr](#):
>
> @samoconnor has done some really nice work reviewing the conventions for `read` / `write` APIs in Base:

This is really interesting as I’ve run into questions about these issues myself. Would it be fair to say that much of the IO behavior in Julia is in flux and expected to have sizeable changes in the near future?

Coming into this I had envisioned something where the IO stream could be handled more like an AbstractArray type with something like getindex for subsetting. I know that there were slightly similar APIs in those packages mentioned but I’d probably need something that preserves aspects of dimensionality while chunking a stream.

---

<div class="post-metadata">

**Author:** ![Zach\_Christensen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zach_christensen/32/7220_2.png) [@Zach\_Christensen](https://discourse.julialang.org/u/Zach_Christensen)\
**Post date:** [January 12, 2019, 4:55pm UTC](https://discourse.julialang.org/t/convention-package-for-chunking-array-stream/19395/5 "2019-01-12T16:55:27Z")

</div>

So after investigating what is out there I put something together.

> **[GitHub - Tokazama/ArrayStreams.jl: Architecture for streaming arrays in Julia.](https://github.com/Tokazama/ArrayStreams.jl)**
>
> Architecture for streaming arrays in Julia. Contribute to Tokazama/ArrayStreams.jl development by creating an account on GitHub.

It’s not ready for use, as I have no tests, package structure, or documentation. The workhorse is `ArrayStream` which is a subtype of AbstractArray. This allows some pretty convenient things when wrapped with an `AxisArray` or `ImageMeta`.
