# Reading files embedded in a Zip-file

**URL:** <https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675>\
**Category:** General Usage\
**Tags:** zip\
**Created:** [May 3, 2018, 9:23am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675 "2018-05-03T09:23:04Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![johann.spies](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johann.spies/32/8805_2.png) [@johann.spies](https://discourse.julialang.org/u/johann.spies)\
**Post date:** [May 3, 2018, 9:23am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/1 "2018-05-03T09:23:04Z")

</div>

I am exploring InfoZIP.

Variable “fn” contains the filename.

```julia
julia> d = open_zip(fn)
julia> d.keys
6-element Array{AbstractString,1}:
 "WR_2004_20180215203106_CORE_0001.xml.gz"
 "WR_2004_20180215203106_CORE_0002.xml.gz"
 "WR_2004_20180215203106_CORE_0003.xml.gz"
 "WR_2004_20180215203106_CORE_0004.xml.gz"
 "Daily_report_CORE_20180215203106.csv"   
 "Y2D_report_CORE_20180215203106.csv"    

```

From there I can do

```julia
julia> l = d["Y2D_report_CORE_20180215203106.csv"];

```

but I cannot do

```julia
julia> using FileIO, CSVFiles, DataFrames
julia> l = load(d["Y2D_report_CORE_20180215203106.csv"]) |> DataFrame
ERROR: stat: name too long (ENAMETOOLONG)
Stacktrace:
 [1] stat(::String) at ./stat.jl:69
 [2] isfile at ./stat.jl:279 [inlined]
 [3] query(::String) at /home/js/.julia/v0.6/FileIO/src/query.jl:377
 [4] #load#13(::Array{Any,1}, ::Function, ::String) at /home/js/.julia/v0.6/FileIO/src/loadsave.jl:52
 [5] load(::String) at /home/js/.julia/v0.6/FileIO/src/loadsave.jl:52

```

Now my question: How can I read those CSV-files into DataFrames and how do I use GZip on the others?

```julia
julia> fh = GZip.open(d["WR_2004_20180215203106_CORE_0001.xml.gz"])

ERROR: MethodError: no method matching gzopen(::Array{UInt8,1})
Closest candidates are:
  gzopen(::AbstractString) at /home/js/.julia/v0.6/GZip/src/GZip.jl:263
  gzopen(::AbstractString, ::AbstractString) at /home/js/.julia/v0.6/GZip/src/GZip.jl:262
  gzopen(::AbstractString, ::AbstractString, ::Integer) at /home/js/.julia/v0.6/GZip/src/GZip.jl:244
  ...
Stacktrace:
 [1] open(::Array{UInt8,1}, ::Vararg{Array{UInt8,1},N} where N) at /home/js/.julia/v0.6/GZip/src/GZip.jl:264

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 3, 2018, 9:31am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/2 "2018-05-03T09:31:30Z")

</div>

It is not very clear what you are trying to do. What are the values of `d`, eg `d["WR_2004_20180215203106_CORE_0001.xml.gz"]`? It looks like byte vector. Why not just use a vector of strings?

FWIW, I would use

> **[GitHub - JuliaIO/CodecZlib.jl: zlib codecs for TranscodingStreams.jl.](https://github.com/JuliaIO/CodecZlib.jl)**
>
> zlib codecs for TranscodingStreams.jl. Contribute to JuliaIO/CodecZlib.jl development by creating an account on GitHub.

to open a stream, and pass it to whatever function I would use otherwise (assuming it accepts streams).

---

<div class="post-metadata">

**Author:** ![johann.spies](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johann.spies/32/8805_2.png) [@johann.spies](https://discourse.julialang.org/u/johann.spies)\
**Post date:** [May 3, 2018, 9:44am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/3 "2018-05-03T09:44:50Z")

</div>

What I want to do is to parse the gzipped XML file(“WR\_2004\_20180215203106\_CORE\_0001.xml.gz”) or put the csv-file in a DataFrame.

Your remark “Why not just use a vector of strings?” is a bit over my head. I have no idea what you mean.

I still do not know how to open it even using CodecZlib.jl:

```julia
julia> using CodecZlib
julia> stream = GzipDecompressorStream(open(d["WR_2004_20180215203106_CORE_0001.xml.gz"]))
ERROR: MethodError: no method matching open(::Array{UInt8,1})
Closest candidates are:
  open(::AbstractString) at iostream.jl:113
  open(::AbstractString, ::Bool, ::Bool, ::Bool, ::Bool, ::Bool) at iostream.jl:103
  open(::AbstractString, ::AbstractString) at iostream.jl:132
  ...

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 3, 2018, 9:49am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/4 "2018-05-03T09:49:38Z")

</div>

1. What tool would you want to use for XML parsing? XML is a generic format.

2. Again, what are the values of `d`? You just listed the keys.

3. I would just use a vector of strings to pass the filenames, eg

---

<div class="post-metadata">

**Author:** ![johann.spies](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johann.spies/32/8805_2.png) [@johann.spies](https://discourse.julialang.org/u/johann.spies)\
**Post date:** [May 3, 2018, 9:56am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/5 "2018-05-03T09:56:22Z")

</div>

Sorry, but I do not now what MWE is.

Here is more information of the zip-archive:

```julia
julia> for (filename, data) in open_zip(fn)
           println("$filename has $(length(data)) bytes and is of type($(typeof(data))")
       end
WR_2004_20180215203106_CORE_0001.xml.gz has 1205519728 bytes and is of type(Array{UInt8,1}
WR_2004_20180215203106_CORE_0002.xml.gz has 1206399676 bytes and is of type(Array{UInt8,1}
WR_2004_20180215203106_CORE_0003.xml.gz has 1203918434 bytes and is of type(Array{UInt8,1}
WR_2004_20180215203106_CORE_0004.xml.gz has 80509285 bytes and is of type(Array{UInt8,1}
Daily_report_CORE_20180215203106.csv has 1499 bytes and is of type(String
Y2D_report_CORE_20180215203106.csv has 1682 bytes and is of type(String

```

I am planning to use LightXML to parse it.

Regards  
Johann

---

<div class="post-metadata">

**Author:** ![johann.spies](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johann.spies/32/8805_2.png) [@johann.spies](https://discourse.julialang.org/u/johann.spies)\
**Post date:** [May 3, 2018, 10:07am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/6 "2018-05-03T10:07:59Z")

</div>

In Python I would do this:

```julia
def openfile(filepath):
    """
    f = file as opened
    t = type of file ('z' for zipfile, 'tgz' for tar.gz, 'gz')
    """
    f = None
    t = None
    if guess_type(filepath) == ('application/x-tar', 'gzip'):
        t = 'tgz'
        f = tarfile.open(filepath,"r:gz")

    elif guess_type(filepath)[1] == 'gzip':
        try:
            f = open(filepath, 'rb')
            t = 'gz'
        except IOError as e:
            print ('Oh dear.')
    elif guess_type(filepath)[0] == 'application/zip':
        try:
            f = zipfile.ZipFile(filepath, 'r')
            t = 'z'
        except IOError as e:
            print ('Oh dear.')

    else:
        try:
            f = open(filepath, 'rb')
            t = None
        except IOError as e:
            print ('Oh dear.')
    return f,t
 f,t = openfile(filename)
if t == 'z':
        gzipfiles = [each for each in f.namelist() if each.endswith('.gz')]
        for gzipf in gzipfiles:
            gzipf_object = f.open(gzipf)
            xml = gzip.GzipFile(fileobj=BytesIO(gzipf_object.read()))
# et cetera

```

I want to get to the point using Julia where I can use the file embedded in a zip-file like this.

---

<div class="post-metadata">

**Author:** ![jandehaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jandehaan/32/6805_2.png) [@jandehaan](https://discourse.julialang.org/u/jandehaan)\
**Post date:** [May 3, 2018, 3:38pm UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/7 "2018-05-03T15:38:37Z")

</div>

MWE: Minimum Working Example

---

<div class="post-metadata">

**Author:** ![js135005](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/js135005/32/8219_2.png) [@js135005](https://discourse.julialang.org/u/js135005)\
**Post date:** [May 3, 2018, 3:56pm UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/8 "2018-05-03T15:56:22Z")

</div>

For what it’s worth, I have found TranscodingStreams together with CodecZlib to be very efficient in handling gzip files with less than 10% slowdown compared to reading the uncompressed files. However for zip archives (.zip), ZipFiles is outdated and is over 20 times slower than processing uncompressed.

I had hoped to either modernize ZipFiles.jl or to write a CodecZlib-like package for zip archives but haven’t been able to make the time to do so. As I result I have dropped support for zip archives and am recommending that our users use gzip compression instead.

---

<div class="post-metadata">

**Author:** ![johann.spies](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johann.spies/32/8805_2.png) [@johann.spies](https://discourse.julialang.org/u/johann.spies)\
**Post date:** [May 7, 2018, 8:01am UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/9 "2018-05-07T08:01:33Z")

</div>

Thanks @jandehaan and @js135005.

My lack of experience in both Julia and using this type of streaming becomes clear.

```julia
julia> using CodecZlib

julia> text = open("2004_CORE.zip")
IOStream(<file 2004_CORE.zip>)

```

From previous experiene (see previous post in this thread) I know there are more than one .gz-file and more than one plain text .csv-file in this zipfile.

Using streaming, how do I find out which files are present, what their names and types are? And then how do I read them seperately or just one of them?

---

<div class="post-metadata">

**Author:** ![js135005](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/js135005/32/8219_2.png) [@js135005](https://discourse.julialang.org/u/js135005)\
**Post date:** [May 8, 2018, 3:28pm UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/10 "2018-05-08T15:28:13Z")

</div>

This is the issue that I’ve hit as well. A zip archive has a directory listing all the included files and the offsets to get to them. The ZipFiles package handles this and allows you to read/write the component files but is abysmally slow so I’ve abandoned using it because I’m handling way too much data.

It desperately needs an update or to be replaced entirely.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [September 2, 2024, 3:46pm UTC](https://discourse.julialang.org/t/reading-files-embedded-in-a-zip-file/10675/11 "2024-09-02T15:46:47Z")

</div>

I created [GitHub - JuliaIO/ZipArchives.jl: Read and write Zip archive files in Julia.](https://github.com/JuliaIO/ZipArchives.jl) which uses CodecZlib.jl internally for decompression and compression. It can efficiently handle reading and writing large amounts of data.
