# Speed comparison on reading a gzip file

**URL:** <https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471>\
**Category:** General Usage\
**Tags:** performance\
**Created:** [March 4, 2021, 11:12am UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471 "2021-03-04T11:12:32Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 4, 2021, 11:12am UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/1 "2021-03-04T11:12:32Z")

</div>

Hi there,

perl using shell piping took ~32 seconds.  
python using shell piping took ~ 36 seconds.  
julia using shell piping took ~ 70 seconds.  
julia using GzipDecompressorStream from CodecZlib, which has been recommended to me by more than one julia user, took ~ 110 seconds.

The gz file is from [https://ftp.ncbi.nlm.nih.gov/gene/DATA/gene2accession.gz](https://ftp.ncbi.nlm.nih.gov/gene/DATA/gene2accession.gz).

Here are the codes.  
perl:

```perl
perl -e 'open( DATA, "pigz -cd ~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz |");
while( <DATA>) { $n = $n + 1}
print( $n, "\n");'

```

python:

```python
import subprocess
with subprocess.Popen( "pigz -cd ~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz", shell = True, stdout = subprocess.PIPE) as gz:
    n1 = 0
    for line1 in gz.stdout:
        n1 = n1 + 1
    print( n1)

```

julia:

```julia
function fun1( file1)
    open( `pigz -cd $( expanduser( file1))`) do io
        n1 = 0
        for line1 in eachline( io)
            n1 = n1 + 1
        end
        print( n1, '\n')
    end
end
fun1( "~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz")

```

julia using GzipDecompressorStream from CodecZlib:

```julia
using CodecZlib, TranscodingStreams
function fun1( file1)
    io1 = GzipDecompressorStream( open( expanduser( file1)))
    n1 = 0
    for line1 in eachline( io1)
        n1 = n1 + 1
    end
    print( n1, '\n')
end
fun1( "~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz")

```

---

<div class="post-metadata">

**Author:** ![jonathanBieler](https://avatars.discourse-cdn.com/v4/letter/j/82dd89/32.png) [@jonathanBieler](https://discourse.julialang.org/u/jonathanBieler)\
**Post date:** [March 4, 2021, 3:32pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/2 "2021-03-04T15:32:21Z")

</div>

Did you run the Julia code twice to make sure you aren’t taking compilation time in your Julia timings ? That said apparently you’ve not the first one to observe this (e.g. see [this discussion](https://github.com/BioJulia/FASTX.jl/issues/40) on reading fastq.gz files).

If there’s really that big of a difference that’s a big opportunity for improvement, since many file type in bioinformatics comes gzip’ed.

---

<div class="post-metadata">

**Author:** ![tim.holy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tim.holy/32/52_2.png) [@tim.holy](https://discourse.julialang.org/u/tim.holy)\
**Post date:** [March 4, 2021, 3:58pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/3 "2021-03-04T15:58:37Z")

</div>

I took a quick profile, and most of the time is spent in a single `ccall`: [CodecZlib.jl/libz.jl at a777d8f53aebd223fe7c7399436a5050784d210f · JuliaIO/CodecZlib.jl · GitHub](https://github.com/JuliaIO/CodecZlib.jl/blob/a777d8f53aebd223fe7c7399436a5050784d210f/src/libz.jl#L93). Interestingly, it’s coming from an `eof` call, but a brief inspection suggested it’s behaving sensibly. So short of rewriting he C code in more optimal form in Julia, it’s nontrivial to know what to do here.

That said, this isn’t really my area so others who need this more should take a look.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 4, 2021, 4:05pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/4 "2021-03-04T16:05:46Z")

</div>

Just now I rerun the function call, and the times were close to previous ones.

---

<div class="post-metadata">

**Author:** ![hannesbecher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hannesbecher/32/28382_2.png) [@hannesbecher](https://discourse.julialang.org/u/hannesbecher)\
**Post date:** [December 16, 2022, 5:56pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/5 "2022-12-16T17:56:22Z")

</div>

I am wondering is there an update on this? I am planning to process large gz-compressed FASTQ files.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [December 16, 2022, 6:08pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/6 "2022-12-16T18:08:12Z")

</div>

You could try [`LibDeflate.jl`](https://github.com/jakobnissen/LibDeflate.jl), which is maintained by a fellow bioinformatician, so probably with needs similar to yours.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [December 21, 2022, 6:49am UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/7 "2022-12-21T06:49:08Z")

</div>

but libdeflate does not support streaming

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [April 17, 2023, 9:20am UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/8 "2023-04-17T09:20:45Z")

</div>

as 2023-4-17, julia version 1.8.5, there is no improvement for julia

---

<div class="post-metadata">

**Author:** ![Xijiang\_Yu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xijiang_yu/32/14439_2.png) [@Xijiang\_Yu](https://discourse.julialang.org/u/Xijiang_Yu)\
**Post date:** [December 10, 2023, 8:45pm UTC](https://discourse.julialang.org/t/speed-comparison-on-reading-a-gzip-file/56471/9 "2023-12-10T20:45:35Z")

</div>

Tried this with Julia 1.9.4:

- Python 35.26s
- Julia-1 1m3s
- Julia-2 1m9s
- `time pigz -dc gene2accession.gz | wc` \to 1m10s
- `time pigz -dc gene2accession.gz | wc -l` \to 34.6s
- `time gzip -dc gene2accession.gz | wc` \to 1m16s
- `time gzip -dc gene2accession.gz | wc -l` \to 1m15s

Seems the bottleneck is `wc`/line counting. Julia is kind of equivalent to `wc`, and python is something like `wc -l`.
