# How to save an array to disk in compressed form?

**URL:** <https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325>\
**Category:** General Usage\
**Tags:** question, data-compression\
**Created:** [May 28, 2020, 8:30am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325 "2020-05-28T08:30:51Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![andrey2185](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andrey2185/32/9889_2.png) [@andrey2185](https://discourse.julialang.org/u/andrey2185)\
**Post date:** [May 28, 2020, 8:30am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/1 "2020-05-28T08:30:51Z")

</div>

Hello, how to save an array to disk in compressed form?

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [May 28, 2020, 10:16am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/2 "2020-05-28T10:16:07Z")

</div>

How about [HDF5](https://github.com/JuliaIO/HDF5.jl) or [FITS](https://juliaastro.github.io/FITSIO.jl/stable/) formats? You may also consider [JLD](https://github.com/JuliaIO/JLD.jl).

---

<div class="post-metadata">

**Author:** ![HenrikM](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrikm/32/17975_2.png) [@HenrikM](https://discourse.julialang.org/u/HenrikM)\
**Post date:** [May 28, 2020, 10:19am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/3 "2020-05-28T10:19:39Z")

</div>

I just asked almost the same question. 🙂 Is [Append to zipped CSV file](https://discourse.julialang.org/t/append-to-zipped-csv-file/40321) of any help? The compression is not much compared to a binary format, though. A rough binary format is [https://docs.julialang.org/en/v1/stdlib/Serialization/](https://docs.julialang.org/en/v1/stdlib/Serialization/). Also search for threads on this site, like [Binary output](https://discourse.julialang.org/t/binary-output/23276), [How store this variable into files, and reaload it?](https://discourse.julialang.org/t/how-store-this-variable-into-files-and-reaload-it/35636), or just search for HDF5 on this site.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 28, 2020, 10:25am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/4 "2020-05-28T10:25:47Z")

</div>

If you store your array in DataFrames format (or any Tables.jl compatible format) then you can use [JDF.jl](https://github.com/xiaodaigh/JDF.jl).

If you don’t need interop with R then Blosc.jl is quite good.

```julia
uncompressed = rand(1_000_000)
using Blosc
compressed = compress(uncompressed)

using Serialization
serialize("somewhere.jls", compressed)

# to read it back
compressed_read_back = deserialize("somewhere.jls")
decompressed = Blosc.decompress(Float64, compressed_read_back)

decompressed == uncompressed # true

```

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [October 31, 2022, 9:28am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/5 "2022-10-31T09:28:17Z")

</div>

See this [github comment](https://github.com/JuliaLang/julia/issues/30133#issuecomment-974692854), which works for generic data not just arrays.  
Before running the code, first run

```julia
using TranscodingStreams, CodecZstd

```

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [October 31, 2022, 10:24am UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/6 "2022-10-31T10:24:50Z")

</div>

Bear in mind, JDIF does not support missing / nothing

> <https://github.com/xiaodaigh/JDF.jl/issues/48>
>
> I got the following error while trying to save some files; I'll upload a MWE lat…er.
> 
> \`\`\`
> ERROR: TaskFailedException:
> MethodError: compress\_then\_write(::Array{Any,1}, ::BufferedStreams.BufferedOutputStream{IOStream}) is ambiguous. Candidates:
> compress\_then\_write(b::Array{Union{Missing, T},1}, io) where T in JDF at /users/yh31/.julia/packages/JDF/jDvZp/src/type-writer-loader/Missing.jl:7
> compress\_then\_write(b::Array{Union{Nothing, T},1}, io) where T in JDF at /users/yh31/.julia/packages/JDF/jDvZp/src/type-writer-loader/Nothing.jl:10
> To resolve the ambiguity, try making one of the methods more specific, or adding a new method more specific than any of the existing applicable methods.
> Stacktrace:
> \[1\] macro expansion at /users/yh31/.julia/packages/JDF/jDvZp/src/savejdf.jl:69 \[inlined\]
> \[2\] (::JDF.var"#49#52"{String,DataFrame,String,Int64})() at ./threadingconstructs.jl:169
> Stacktrace:
> \[1\] wait at ./task.jl:267 \[inlined\]
> \[2\] fetch(::Task) at ./task.jl:282
> \[3\] \_broadcast\_getindex\_evalf at ./broadcast.jl:648 \[inlined\]
> \[4\] \_broadcast\_getindex at ./broadcast.jl:621 \[inlined\]
> \[5\] getindex at ./broadcast.jl:575 \[inlined\]
> \[6\] copyto\_nonleaf!(::Array{NamedTuple,1}, ::Base.Broadcast.Broadcasted{Base.Broadcast.DefaultArrayStyle{1},Tuple{Base.OneTo{Int64}},typeof(fetch),Tuple{Base.Broadcast.Extruded{Array{Any,1},Tuple{Bool},Tuple{Int64}}}}, ::Base.OneTo{Int64}, ::Int64, ::Int64) at ./broadcast.jl:1026
> \[7\] restart\_copyto\_nonleaf!(::Array{NamedTuple,1}, ::Array{NamedTuple{(:string\_compressed\_bytes, :string\_len\_bytes, :rle\_bytes, :rle\_len, :type, :len),Tuple{Int64,Int64,Int64,Int64,DataType,Int64}},1}, ::Base.Broadcast.Broadcasted{Base.Broadcast.DefaultArrayStyle{1},Tuple{Base.OneTo{Int64}},typeof(fetch),Tuple{Base.Broadcast.Extruded{Array{Any,1},Tuple{Bool},Tuple{Int64}}}}, ::NamedTuple{(:len, :type),Tuple{Int64,DataType}}, ::Int64, ::Base.OneTo{Int64}, ::Int64, ::Int64) at ./broadcast.jl:1017
> \[8\] copyto\_nonleaf!(::Array{NamedTuple{(:string\_compressed\_bytes, :string\_len\_bytes, :rle\_bytes, :rle\_len, :type, :len),Tuple{Int64,Int64,Int64,Int64,DataType,Int64}},1}, ::Base.Broadcast.Broadcasted{Base.Broadcast.DefaultArrayStyle{1},Tuple{Base.OneTo{Int64}},typeof(fetch),Tuple{Base.Broadcast.Extruded{Array{Any,1},Tuple{Bool},Tuple{Int64}}}}, ::Base.OneTo{Int64}, ::Int64, ::Int64) at ./broadcast.jl:1033
> \[9\] copy at ./broadcast.jl:880 \[inlined\]
> \[10\] materialize at ./broadcast.jl:837 \[inlined\]
> \[11\] savejdf(::String, ::DataFrame; verbose::Bool) at /users/yh31/.julia/packages/JDF/jDvZp/src/savejdf.jl:76
> \[12\] savejdf(::String, ::DataFrame) at /users/yh31/.julia/packages/JDF/jDvZp/src/savejdf.jl:48
> \[13\] (::var"#51#52")(::File) at ./REPL\[18\]:2
> \[14\] (::FileTrees.var"#saver#89"{var"#51#52"})(::File, ::DataFrame) at /users/yh31/.julia/packages/FileTrees/sx5xd/src/values.jl:121
> \[15\] (::Dagger.var"#47#48"{FileTrees.var"#saver#89"{var"#51#52"},Tuple{File,DataFrame}})() at ./threadingconstructs.jl:169
> wait at ./task.jl:267 \[inlined\]
> fetch at ./task.jl:282 \[inlined\]
> execute!(::Dagger.ThreadProc, ::Function, ::File, ::Vararg{Any,N} where N) at /users/yh31/.julia/packages/Dagger/U857J/src/processor.jl:222
> do\_task(::Dagger.Context, ::Dagger.OSProc, ::Int64, ::Function, ::Tuple{File,Dagger.Chunk{Any,MemPool.DRef,Dagger.ThreadProc}}, ::Bool, ::Bool, ::Bool, ::Dagger.Sch.ThunkOptions) at /users/yh31/.julia/packages/Dagger/U857J/src/scheduler.jl:340
> \#137 at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/remotecall.jl:354 \[inlined\]
> run\_work\_thunk(::Distributed.var"#137#138"{typeof(Dagger.Sch.do\_task),Tuple{Dagger.Context,Dagger.OSProc,Int64,FileTrees.var"#saver#89"{var"#51#52"},Tuple{File,Dagger.Chunk{Any,MemPool.DRef,Dagger.ThreadProc}},Bool,Bool,Bool,Dagger.Sch.ThunkOptions},Base.Iterators.Pairs{Union{},Union{},Tuple{},NamedTuple{(),Tuple{}}}}, ::Bool) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/process\_messages.jl:79
> remotecall\_fetch(::Function, ::Distributed.LocalProcess, ::Dagger.Context, ::Vararg{Any,N} where N; kwargs::Base.Iterators.Pairs{Union{},Union{},Tuple{},NamedTuple{(),Tuple{}}}) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/remotecall.jl:379
> remotecall\_fetch(::Function, ::Distributed.LocalProcess, ::Dagger.Context, ::Vararg{Any,N} where N) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/remotecall.jl:379
> remotecall\_fetch(::Function, ::Int64, ::Dagger.Context, ::Vararg{Any,N} where N; kwargs::Base.Iterators.Pairs{Union{},Union{},Tuple{},NamedTuple{(),Tuple{}}}) at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/remotecall.jl:421
> remotecall\_fetch at /buildworker/worker/package\_linux64/build/usr/share/julia/stdlib/v1.5/Distributed/src/remotecall.jl:421 \[inlined\]
> macro expansion at /users/yh31/.julia/packages/Dagger/U857J/src/scheduler.jl:353 \[inlined\]
> (::Dagger.Sch.var"#26#27"{Dagger.Context,Dagger.OSProc,Int64,FileTrees.var"#saver#89"{var"#51#52"},Tuple{File,Dagger.Chunk{Any,MemPool.DRef,Dagger.ThreadProc}},Channel{Any},Bool,Bool,Bool,Dagger.Sch.ThunkOptions})() at ./task.jl:356
> Stacktrace:
> \[1\] compute\_dag(::Dagger.Context, ::Dagger.Thunk; options::Nothing) at /users/yh31/.julia/packages/Dagger/U857J/src/scheduler.jl:137
> \[2\] compute(::Dagger.Context, ::Dagger.Thunk; options::Nothing) at /users/yh31/.julia/packages/Dagger/U857J/src/compute.jl:32
> \[3\] #compute#70 at /users/yh31/.julia/packages/Dagger/U857J/src/compute.jl:5 \[inlined\]
> \[4\] compute at /users/yh31/.julia/packages/Dagger/U857J/src/compute.jl:5 \[inlined\]
> \[5\] exec(::Dagger.Thunk) at /users/yh31/.julia/packages/FileTrees/sx5xd/src/parallelism.jl:68
> \[6\] save(::var"#51#52", ::FileTree; lazy::Nothing, exec::Bool) at /users/yh31/.julia/packages/FileTrees/sx5xd/src/values.jl:128
> \[7\] save(::Function, ::FileTree) at /users/yh31/.julia/packages/FileTrees/sx5xd/src/values.jl:111
> \[8\] top-level scope at REPL\[18\]:1
> \`\`\`

---

<div class="post-metadata">

**Author:** ![maxchendt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxchendt/32/42976_2.png) [@maxchendt](https://discourse.julialang.org/u/maxchendt)\
**Post date:** [January 23, 2023, 12:44pm UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/7 "2023-01-23T12:44:52Z")

</div>

I find that using `Serialization + TranscodingStreams + CodecXz` may be a good way, all packages needed are small overhead

```julia
using Downloads, TranscodingStreams, CodecXz, Serialization

xzFile = Downloads.download("https://github.com/PharosAbad/PharosAbad.github.io/raw/master/files/sp500.jls.xz")

io = open(xzFile)
io = TranscodingStream(XzDecompressor(), io)
E = deserialize(io)
V = deserialize(io)
close(io)

```

now, we compress the data

```julia
xzFile = "/tmp/my-sp500.jls.xz"
io = open(xzFile, "w")
io = TranscodingStream(XzCompressor(), io)
serialize(io, E)
serialize(io, V)
close(io)

```

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [January 23, 2023, 3:36pm UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/8 "2023-01-23T15:36:18Z")

</div>

> [@xiaodai](#):
>
> ```julia
> uncompressed = rand(1_000_000)
> using Blosc
> 
> ```

Blosc is a good compressor, one of the best _lossless_ compressor, but you can do much better with a **lossy** compressor. [Uniformly] random info shouldn’t compress at all, `rand` is however for normal distributed, so will compress, and real-world data even more.

For some lossy compressor that’s available with Julia:

> **[GitHub - milankl/ZfpCompression.jl: Julia bindings for the data compression...](https://github.com/milankl/ZfpCompression.jl)**
>
> Julia bindings for the data compression library zfp - GitHub - milankl/ZfpCompression.jl: Julia bindings for the data compression library zfp

> julia\> using ZfpCompression  
> julia\> A = rand(Float32,100,50);  
> julia\> Ac = zfp\_compress(A)

Note right there using Float32, starting from half the size (or possibly even Float16), is useful, for lossy or not, since often the extra 32-bit might not be valuable data. Note, also “reversible (lossless) compression is supported.” And you may want to (compress and) decompress to Float64, for further processing, and there are useful tuning options (for lossy).

I didn’t immediately find Julia software for state-of-the-art SZ3 (or SZ):

[https://szcompressor.org/](https://szcompressor.org/)

or well until (seemingly though only very domain specific):

> **[Compressing multidimensional weather and climate data into neural networks](https://arxiv.org/abs/2210.12538)**
>
> Weather and climate simulations produce petabytes of high-resolution data that are later analyzed by researchers in order to understand climate change or severe weather. We propose a new method of compressing this multidimensional weather and climate...

> While compression ratios range from 300x to more than 3,000x, our method outperforms the state-of-the-art compressor SZ3 in terms of weighted RMSE, MAE.

---

<div class="post-metadata">

**Author:** ![maxchendt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxchendt/32/42976_2.png) [@maxchendt](https://discourse.julialang.org/u/maxchendt)\
**Post date:** [January 23, 2023, 10:45pm UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/9 "2023-01-23T22:45:06Z")

</div>

> [@Palli](#):
>
> since often the extra 32-bit might not be valuable data

That maybe true. But users may have other reasons. As for me, I compressed data because I am using `BigFloat` to maintain precision, and the raw size is 16G+ usually.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [January 24, 2023, 2:32pm UTC](https://discourse.julialang.org/t/how-to-save-an-array-to-disk-in-compressed-form/40325/10 "2023-01-24T14:32:14Z")

</div>

I feel like BigFloat is almost always a mistake, it’s by default 256 bits of precision (you can set it higher, or lower, and I find it likely that that flexibility, even if unused, makes it slower), so right there you use 8x more memory _before compression_ (or 4x compared to Float64), and it’s also a performance killer.

And for what? I would at least consider the much faster (fastest alternative to Float64, with more bits) intermediate Float128 from [GitHub - JuliaMath/Quadmath.jl: Float128 and libquadmath for the Julia language](https://github.com/JuliaMath/Quadmath.jl)

In either case, no matter how many bits you use, you are not immune from catastrophic cancellation, i.e. loss of precision, so I would also consider (ValidatedNumerics.jl and/or):

> **[GitHub - JuliaIntervals/IntervalArithmetic.jl: Library for validated numerics...](https://github.com/JuliaIntervals/IntervalArithmetic.jl)**
>
> Library for validated numerics using interval arithmetic - GitHub - JuliaIntervals/IntervalArithmetic.jl: Library for validated numerics using interval arithmetic

> The final result is an interval that is _guaranteed_ to contain the correct result, starting from the given initial data.

It’s based on Float64 by default, but supports down to Float16, so each number is 2x64 down to 2x16 = 32 bits. Always slower than the types it’s built on (2x slower?), in all cases likely much faster then BigFloat, and more accurate. I’m not sure, but it might transparently support compression, I believe Bloch should work, ZfpCompression might in case it sees e.g. Float64 from inside the struct, but note then to use it’s lossless mode (you could try the lossy mode, but you would likely destroy the guarantee the package gives you, though you might not be far off, and with some support it might be kept).

I see recent:

> ### Breaking changes
> 
> - Changed from using `FastRounding.jl` to `RoundingEmulator.jl` for the [default [fixed typo]] rounding mode. [#370](https://github.com/JuliaIntervals/IntervalArithmetic.jl/pull/370)

Note one other option:

> **[GitHub - milankl/StochasticRounding.jl: Up or down? Maybe both?](https://github.com/milankl/StochasticRounding.jl)**
>
> Up or down? Maybe both? Contribute to milankl/StochasticRounding.jl development by creating an account on GitHub.

> As `1/3` is not exactly representable the rounding will be at 66.6% chance towards 0.33398438 and at 33.3% towards 0.33203125 such that in expectation the result is 0.33333… and therefore exact.

E.g. 1/3 (and 1/10) isn’t exact in any binary floating point, why people may be tempted to use ever higher precision to better approximate, but there’s always an error, and it can grow, but good to know with with the above, even this is very viable, `BFloat16sr` based below, only 63% slower than Float64 with its inferior rounding:

> **[GitHub - JuliaMath/BFloat16s.jl: Nobody needed all those bits anyway](https://github.com/JuliaMath/BFloat16s.jl)**
>
> Nobody needed all those bits anyway. Contribute to JuliaMath/BFloat16s.jl development by creating an account on GitHub.

See also (I just discovered this one): [GitHub - AnderGray/IntervalUnionArithmetic.jl: An implementation of interval union arithmetic in Julia](https://github.com/AnderGray/IntervalUnionArithmetic.jl)
