# \[ANN\] DiskArrays.jl

**URL:** <https://discourse.julialang.org/t/ann-diskarrays-jl/34119>\
**Category:** Package Announcements\
**Created:** [February 3, 2020, 3:14pm UTC](https://discourse.julialang.org/t/ann-diskarrays-jl/34119 "2020-02-03T15:14:29Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![fabiangans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabiangans/32/2624_2.png) [@fabiangans](https://discourse.julialang.org/u/fabiangans)\
**Post date:** [February 3, 2020, 3:14pm UTC](https://discourse.julialang.org/t/ann-diskarrays-jl/34119/1 "2020-02-03T15:14:29Z")

</div>

As already discussed [here](https://discourse.julialang.org/t/taking-the-array-indexing-interface-seriously/32035) I worked on a package that would make it easier to use `AbstractArray`s that map to chunked data on disk.

The package provides an `AbstractDiskArray` which you can subtype to get efficient indexing for common access patterns, efficient reductions and lazy broadcasting for free. Currently this is only implemented for Zarr.jl, but PRs for [NetCDF.jl](https://github.com/JuliaGeo/NetCDF.jl/pull/111) and [ArchGDAL.jl](https://github.com/yeesian/ArchGDAL.jl/pull/105) are pending and other are [planned](https://github.com/meggart/DiskArrays.jl/issues/2).

## Example

We open a Zarr array on GCS. Note that accessing a single element from the array takes about 18s on my machine (limited by internet connection), because a whole chunk of data has to be transferred and decompressed to access the single value. So we open a reference to the Array and use normal boradcasting syntax to convert to Celsius and multiply with latitudinal weights:

```julia
using Zarr, Statistics
g = zopen("gs://cmip6/CMIP/NCAR/CESM2/historical/r9i1p1f1/Amon/tas/gn/", consolidated=true)

latweights = reshape(cosd.(g["lat"])[:],1,192,1)
t_celsius = g["tas"].-273.15
t_w = t_celsius .* latweights

```

```julia-auto
Disk Array with size 288 x 192 x 1980

```

Nothing happens yet and the subsequent broadcast operations will be fused. Then we compute the mean over the spatial dimensions:

```julia
mean(t_w, dims = (1,2))./mean(latweights)

```

which will iterate over each chunk of the array at a time, making sure that every chunk is accessed and transferred only once. So, memory usage is in the order of the size of a single chunk and the whole computation takes about a minute, which equals the time to read the whole array.

## Implementation

I tried to re-use as much of Julia Base functionality as possible to implement the custom reductions and broadcasting. So basically every function that falls back to `mapreduce` or `mapfoldl` (which includes, `maximum`, `minimum`, `any`, `all`, `sum`, …) will have an efficient access pattern.

For broadcasting I just wrap the broadcasted expression into a new `BroadcastDiskArray` type which again subtypes `AbstractDiskArray` to benefit from the changed getindex and mapreduce behavior.

I have to admit it was a tough experience to hook into the mapreduce and broadcast code, but in the end I was surprised to get the behavior described above with only [~140 LOC](https://github.com/meggart/DiskArrays.jl/blob/master/src/ops.jl).

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [February 3, 2020, 6:22pm UTC](https://discourse.julialang.org/t/ann-diskarrays-jl/34119/2 "2020-02-03T18:22:58Z")

</div>

Sounds very interesting!  
Perhaps (as usual) a lunatic comment from me. Tape is of course still heavily used for long term storage.  
I doubt anyone actually reads directly from tapes any more - the data will be staged to disk.  
However I note that LTFS exists

> **[Linear Tape File System](https://en.wikipedia.org/wiki/Linear_Tape_File_System)**
>
> The Linear Tape File System (LTFS) is a file system that allows files stored on magnetic tape to be accessed in a similar fashion to those on disk or removable flash drives. It requires both a specific format of data on the tape media and software to provide a file system interface to the data.
> The technology, based around a self-describing tape format developed by IBM, was adopted by the LTO Consortium in 2010.
> Magnetic tape data storage has been used for over 50 years, but typically did not ...

Would it be lunatic to think that Julia could interact with LTFS?

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 31, 2020, 12:15am UTC](https://discourse.julialang.org/t/ann-diskarrays-jl/34119/3 "2020-10-31T00:15:39Z")

</div>

Do we have a disk data frame library?

---

<div class="post-metadata">

**Author:** ![fabiangans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabiangans/32/2624_2.png) [@fabiangans](https://discourse.julialang.org/u/fabiangans)\
**Post date:** [November 2, 2020, 10:41am UTC](https://discourse.julialang.org/t/ann-diskarrays-jl/34119/4 "2020-11-02T10:41:39Z")

</div>

I don’t know, and I am not sure we need one. I think the [Tables.jl](https://juliadata.github.io/Tables.jl/stable/) interface is carefully designed with out-of-core Tables in mind, and many different data formats supporting table-like data already implement this interface. Where exactly do you see the gap?
