# Compressed array-on-disk

**URL:** <https://discourse.julialang.org/t/compressed-array-on-disk/69410>\
**Category:** Data\
**Tags:** data-compression\
**Created:** [October 8, 2021, 3:58am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410 "2021-10-08T03:58:49Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [October 8, 2021, 3:58am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/1 "2021-10-08T03:58:49Z")

</div>

I am working on an algorithm that fits an HMM model to some data. As part of the fitting I need to run the so-called [forward-backward] algorithm ([Forward–backward algorithm - Wikipedia](https://en.wikipedia.org/wiki/Forward%E2%80%93backward_algorithm)). This involves constructing matrices that easily exceed the available memory and so I am trying to find an efficient way to temporarily write and read these matrices to disk. Because of the way the transition matrices are structured, a lot of the values in these matrices will be identical, and so would probably benefit greatly from some form of compression (.e.g [Blosc.jl](https://github.com/JuliaIO/Blosc.jl)). Ideally, I would like to be able to access these matrices as though they were normal matrices, and have the underlying chunking/compression/decompression handled automatically. I could put this together myself, I think, but I wanted to check if anyone is aware of a package that already does something like this?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 8, 2021, 4:01am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/2 "2021-10-08T04:01:57Z")

</div>

I think memory mapping might do something that you’ve asked for.

Others I have seen, but not necessarily have everything in place but might be worth looking at are Zarr.jl and [DiskArrays.jl](https://github.com/meggart/DiskArrays.jl) might be useful.

If you blosc compress something then you can’t access the elements without uncompress, so you might need to compress in small blocks etc.

Anyway, I also want to know the answer to this but I am just sharing some fragments of things I remember from the past for this.

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [October 8, 2021, 4:10am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/3 "2021-10-08T04:10:01Z")

</div>

Thanks! Those look like valuable resources. The main functionality would be the ability of seamlessly read and write the chunked data. Basically, the chunks would be as large as can fit in memory. I had very manual python version of this at some point where I would write individual chunks to separate files, and then have a bunch bookkeeping code to keep track of the chunks and loading the correct one. This worked, but made the code a bit messy, so I was hoping to refactor that part out into a separate package in my julia implementation. I will look into Zarr and DiskArrays in more details and see if they fulfil my needs.

---

<div class="post-metadata">

**Author:** ![laborg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/laborg/32/5474_2.png) [@laborg](https://discourse.julialang.org/u/laborg)\
**Post date:** [October 8, 2021, 5:25am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/4 "2021-10-08T05:25:33Z")

</div>

I would use an uncompressed HDF5 with mmap. This should be quite efficient in terms of LOC on your side.

---

<div class="post-metadata">

**Author:** ![fabiangans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabiangans/32/2624_2.png) [@fabiangans](https://discourse.julialang.org/u/fabiangans)\
**Post date:** [October 8, 2021, 6:26am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/5 "2021-10-08T06:26:02Z")

</div>

Zarr.jl might also be an option for you:

[https://github.com/meggart/Zarr.jl](https://github.com/meggart/Zarr.jl)

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [October 8, 2021, 8:06am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/6 "2021-10-08T08:06:15Z")

</div>

Am I correct in my understanding that Zarr arrays are fast if you read them in chunks? In other words, if I try to access single elements of a Zarr array in my code, for example in a loop, that would be slow?

---

<div class="post-metadata">

**Author:** ![fabiangans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabiangans/32/2624_2.png) [@fabiangans](https://discourse.julialang.org/u/fabiangans)\
**Post date:** [October 8, 2021, 8:11am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/7 "2021-10-08T08:11:41Z")

</div>

Yes, Zarr is slow for random access of single data points, but this will be tru for every compressed data format.

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [October 8, 2021, 8:15am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/8 "2021-10-08T08:15:38Z")

</div>

Thanks for the clarification. What I had in mind is something that automatically figures out which chunk the requested index “belongs to” and then reads the chunk of data and decompresses it in memory. This would work will for my usage case, where I’m accessing the data sequentially, i.e. column by column. All I would need to do is to check when the index “crosses the chunk border” and the load the next chunk.  
It sounds like I could probably add this functionality on top of Zarr, again for my very special usage case.

---

<div class="post-metadata">

**Author:** ![fabiangans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabiangans/32/2624_2.png) [@fabiangans](https://discourse.julialang.org/u/fabiangans)\
**Post date:** [October 11, 2021, 11:27am UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/9 "2021-10-11T11:27:10Z")

</div>

> [@grero](#):
>
> It sounds like I could probably add this functionality on top of Zarr, again for my very special usage case.

I am not sure if your use case is so special, to me it actually sounds quite common. In YAXArrays.jl we have implemented a chunk-aware `mapslices`, which should work out-of-the box if your input data comes from NetCDF or xarray-compatible Zarr. You just map your function over your data and in the background the data is read chunk by chunk, processed and written to disk again. You also get multi-threading and distributed-processing for free. Unfortunately the documentation has not yet been migrated from ESDL.jl, so if you want to read about it, you could start here: [https://esa-esdl.github.io/ESDL.jl/latest/analysis/#Applying-custom-functions-1](https://esa-esdl.github.io/ESDL.jl/latest/analysis/#Applying-custom-functions-1)

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [October 11, 2021, 1:36pm UTC](https://discourse.julialang.org/t/compressed-array-on-disk/69410/10 "2021-10-11T13:36:51Z")

</div>

That sounds like exactly what I need, I’ll check it out. Thanks!
