# How to unpack an .xz file with Julia

**URL:** <https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028>\
**Category:** General Usage\
**Tags:** question\
**Created:** [April 26, 2021, 11:53am UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028 "2021-04-26T11:53:43Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 26, 2021, 11:53am UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/1 "2021-04-26T11:53:44Z")

</div>

Is there a package that can decompress .xz files?

I need to do this in a cross-platform way.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [April 26, 2021, 12:07pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/2 "2021-04-26T12:07:54Z")

</div>

[https://github.com/JuliaIO/TranscodingStreams.jl](https://github.com/JuliaIO/TranscodingStreams.jl) + [https://github.com/JuliaIO/CodecXz.jl](https://github.com/JuliaIO/CodecXz.jl)

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 26, 2021, 12:13pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/3 "2021-04-26T12:13:46Z")

</div>

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 26, 2021, 3:25pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/4 "2021-04-26T15:25:42Z")

</div>

The following code works for me:

```julia
using CodecXz

const FILENAME="data/log_8700W_8ms.csv.xz"

stream = open(FILENAME)
output = open(FILENAME[1:end-3],"w")
for line in eachline(XzDecompressorStream(stream))
    println(output, line)
end
close(stream)
close(output)

```

But it will work only for decompressing text files.

How can it be generalized for any files?

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [April 26, 2021, 3:32pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/5 "2021-04-26T15:32:18Z")

</div>

What do you mean by “any file”? Code you wrote is generic enough, there should be no difference in decompressing text or any other file (in the end of the day all files are just text).

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [April 26, 2021, 4:04pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/6 "2021-04-26T16:04:35Z")

</div>

I mean, eachline will only work if the stream contains line delimiters, or am I wrong?

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [April 26, 2021, 4:53pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/7 "2021-04-26T16:53:25Z")

</div>

Ah, that’s the best part actually! Since TranscodingStream produces `IO` object, all [General IO](https://docs.julialang.org/en/v1/base/io-network/#General-I/O) applies. It doesn’t matter whether it is compressed data or not at this point.

All what follows depends on packages that you use or procedure that you need to implement. You can materialize data as `Vector{UInt8}` or any other data format. Or maybe your package can accept this `IO` object and you can forget about compressed data processing completely.

As an example, consider following xz arrow manipulations

```julia
using CodecXz
using Arrow
using Tables

x = [(; a = 1, b = 2)]
Arrow.write("x.arrow", x)

```

Here we switch to shell and compress data manually. It can be done with the `CodecXz` of course, but we pretend that this is external file.

```bash
sh> xz x.arrow

```

and back to Julia

```julia
julia> stream = XzDecompressorStream(open("x.arrow.xz"))
TranscodingStreams.TranscodingStream{XzDecompressor, IOStream}(<mode=idle>)

# we can materialize uncompressed data as Vector{UInt8}
julia> read(stream)
610-element Vector{UInt8}:
 0x41
 0x52
 0x52
    ⋮

# we can read it as a String (which is weird of course, since it is binary file)
julia> read(stream, String)
"ARROW1\0\0\xff\xff\xff\xff\xa8\0\0\0\x10\0\0\0\0\0\n\0\f\0\n\0\b\0\x04\0\n\0\0\0\x10\0\0\0\x01\
0\x04\0\b\0\b\0\0\0\x04\0\b\0\0\0\x04\0\0\0\x02\0\0\0D\0\0\0\x04\0\0\0\xd4\xff\xff\xff\x10\0\0\0..."

# and we can read it back as Julia data
julia> Arrow.Table(stream) |> Tables.rowtable
1-element Vector{NamedTuple{(:a, :b), Tuple{Int64, Int64}}}:
 (a = 1, b = 2)

```

One small note: in these manipulations, before each `read` I actually use `stream = XzDecompressorStream(open("x.arrow.xz"))` command, because you can’t `read` twice from the same stream. I just omit it for simplicity.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 26, 2021, 5:14pm UTC](https://discourse.julialang.org/t/how-to-unpack-an-xz-file-with-julia/60028/8 "2021-04-26T17:14:37Z")

</div>

The only functions in your snippet that care about line delimiters are `eachline` and `println`. For example, `read(XzDecompressorStream(stream))` would read the whole decompressed content as a vector of raw bytes.
