# Slow gzip streaming in julia but not in python

**URL:** <https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346>\
**Category:** General Usage\
**Tags:** performance\
**Created:** [March 2, 2021, 4:08pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346 "2021-03-02T16:08:44Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 4:08pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/1 "2021-03-02T16:08:44Z")

</div>

Hi there,

I am new to Julia, and come to Julia for performance with fewer codes. I have been using R, and encountered a simple task: read from a gziped file and do something to each line. It was slow in R, of course. So I went to Perl first. Here is the code:

```
perl -e '
open( DATA, "zcat ~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz|");
while( <DATA>) {
  @F = split( /\t/);
}
close( DATA);'

```

I found the gzip command was using about 48% CPU of a thread. Then I came to Julia, and here is the code:

```
function fun1( file1)
    for line1 in eachline( `zcat $( expanduser( file1))`)
        F = split( line1, '\t')
    end
end
fun1( "~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz")

```

And I fount that the gzip command was also using about 48% CPU. Then I came to Python, and here is the code:

```
import subprocess
with subprocess.Popen( "zcat ~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz", shell = True, stdout = subprocess.PIPE) as gz:
    for line1 in gz.stdout:
        F = line1.split( b'\t')

```

Now, the gzip is using nearly 100% CPU, and so is Python.

I have not recoreded times spent by these 3 approaches. But I can feel that the python code is obviously faster than the other two.

How come?

---

<div class="post-metadata">

**Author:** ![simonbyrne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simonbyrne/32/19_2.png) [@simonbyrne](https://discourse.julialang.org/u/simonbyrne)\
**Post date:** [March 2, 2021, 4:21pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/2 "2021-03-02T16:21:56Z")

</div>

How does the timed performance compare?

I’m not sure it will make much difference in this case, but it is generally [bad for performance](https://docs.julialang.org/en/v1/manual/performance-tips/#Avoid-global-variables) to refer to non-`const` global variables inside a function.

Also, I would recommend using [CodecZlib.jl](https://github.com/JuliaIO/CodecZlib.jl) over shell commands.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 4:28pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/3 "2021-03-02T16:28:20Z")

</div>

I will later record the times spent. But I have the impression that the python code is obviously faster than the other two.

For the julia code, there is only one global variable, which was not used repeatedly. So I guess it would not affect the performance much.

---

<div class="post-metadata">

**Author:** ![WschW](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wschw/32/6575_2.png) [@WschW](https://discourse.julialang.org/u/WschW)\
**Post date:** [March 2, 2021, 4:35pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/4 "2021-03-02T16:35:41Z")

</div>

The issue with non const global variables is in how they affect type inference. Basically it and every value derived from that are a type which cannot be known at compile time which results in slower code.

If there are still speed differences after that it could be due to whether `println` buffers IO operations.

---

<div class="post-metadata">

**Author:** ![rdeits](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rdeits/32/286_2.png) [@rdeits](https://discourse.julialang.org/u/rdeits)\
**Post date:** [March 2, 2021, 4:40pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/5 "2021-03-02T16:40:23Z")

</div>

This could also be related to the fact that `println` is slower than Python’s `print` in some terminals: [https://github.com/JuliaLang/julia/issues/36639](https://github.com/JuliaLang/julia/issues/36639)

Do your results change if you skip the `print` statements?

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 4:46pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/6 "2021-03-02T16:46:31Z")

</div>

I see. I have modified the code, but noticed no improvement in performance.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 4:48pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/7 "2021-03-02T16:48:19Z")

</div>

I later skipped printing, but noticed no improvement in performance - gzip was using ~45% CPU, and Julia 100%.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 2, 2021, 4:48pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/8 "2021-03-02T16:48:29Z")

</div>

Yea why are you calling externally here and not using a Julia package? I read in large `*.csv.gz` files via streaming using CSV.jl and CodecZLib.jl and it is very fast.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 5:36pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/9 "2021-03-02T17:36:52Z")

</div>

Is this code appropriate in using CodecZlib? It is even more slow…

```
using CSV, CodecZlib, Mmap, TranscodingStreams
function fun1( file1)
    io1 = TranscodingStream( GzipDecompressor(), open( expanduser( file1)))
    N = 10
    n1 = 0
    for line1 in eachline( io1)
        F = split( line1, '\t')
        if n1 < N # print the first 10 lines
            print( line1, '\n')
            n1 = n1 + 1
        end
    end
end
fun1( "~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz")

```

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [March 2, 2021, 6:32pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/10 "2021-03-02T18:32:59Z")

</div>

> [@calvin](#):
>
> ` "zcat ~/in2/swxx/sj/db/ncbi/gene/gene2accession.gz|"`

just a side note - with a multi core cpu :  
the "`unpigz -c sample.gz | .... `" is ~ 2x faster than a simple `zcat sample.gz | ... `  
[https://unix.stackexchange.com/questions/363644/fastest-and-most-efficient-way-to-get-number-of-records-lines-in-a-gzip-compre/363739#363739](https://unix.stackexchange.com/questions/363644/fastest-and-most-efficient-way-to-get-number-of-records-lines-in-a-gzip-compre/363739#363739)

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 2, 2021, 10:48pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/11 "2021-03-02T22:48:54Z")

</div>

I know. However, gzip is not the bottleneck here.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [March 3, 2021, 12:15am UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/12 "2021-03-03T00:15:53Z")

</div>

There was a [related thread](https://discourse.julialang.org/t/bad-performance-of-eachline-on-stdin/55701) recently where it was mentioned that reading stdin via a pipe can be quite slow in Julia. `CodecZlib` definitely seems like the more idiomatic option (just like you’d use `gzip.open` in Python), so hopefully someone can comment on why the performance of your code snippet isn’t amazing (lack of buffering, maybe)?

---

<div class="post-metadata">

**Author:** ![simonbyrne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simonbyrne/32/19_2.png) [@simonbyrne](https://discourse.julialang.org/u/simonbyrne)\
**Post date:** [March 3, 2021, 9:08pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/13 "2021-03-03T21:08:06Z")

</div>

I think that’s right (you should be able to just do `GzipDecompressorStream(open( expanduser( file1)))`, but that is equivalent).

Is the file (or something similar) publicly available? If you can post a link, I suspect you will find some people who will try to optimise it.

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 4, 2021, 1:28am UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/14 "2021-03-04T01:28:45Z")

</div>

OK. Thanks for the answer. The file is 2+ gb and is publicly available: [https://ftp.ncbi.nlm.nih.gov/gene/DATA/gene2accession.gz](https://ftp.ncbi.nlm.nih.gov/gene/DATA/gene2accession.gz).

---

<div class="post-metadata">

**Author:** ![simonbyrne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simonbyrne/32/19_2.png) [@simonbyrne](https://discourse.julialang.org/u/simonbyrne)\
**Post date:** [March 4, 2021, 7:18pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/15 "2021-03-04T19:18:58Z")

</div>

A post was split to a new topic: [Edit limits](https://discourse.julialang.org/t/edit-limits/56495)

---

<div class="post-metadata">

**Author:** ![simonbyrne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simonbyrne/32/19_2.png) [@simonbyrne](https://discourse.julialang.org/u/simonbyrne)\
**Post date:** [March 4, 2021, 8:31pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/16 "2021-03-04T20:31:52Z")

</div>

Thanks, doing some basic timings, I think there are actually two distinct performance problems:

1. calling `eachline`: simply iterating over the `gzcat` pipe takes ~1 minute in julia, vs ~30 seconds in Python. CodecZlib.jl seems to have similar performance to `gzcat`.
2. `split` in Julia is itself slower than in Python

Not sure exactly what the best approach here would be.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 17, 2021, 1:07pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/17 "2021-03-17T13:07:17Z")

</div>

Out of curiosity, you could try LibDeflate.jl: [[ANN] LibDeflate.jl and CodecBGZF.jl - really fast blocked de/compression - #8 by jakobnissen](https://discourse.julialang.org/t/ann-libdeflate-jl-and-codecbgzf-jl-really-fast-blocked-de-compression/45410/8)

---

<div class="post-metadata">

**Author:** ![calvin](https://avatars.discourse-cdn.com/v4/letter/c/f475e1/32.png) [@calvin](https://discourse.julialang.org/u/calvin)\
**Post date:** [March 17, 2021, 2:08pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/18 "2021-03-17T14:08:07Z")

</div>

However, streaming is not supported…

> Unlike libz or gzip, libdeflate does not support streaming, and so is intended for use in of files that fit in-memory or for block-compressed files like bgzip.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 17, 2021, 2:08pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/19 "2021-03-17T14:08:56Z")

</div>

Ah, sorry, I’ve missed it.

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [March 17, 2021, 3:20pm UTC](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346/20 "2021-03-17T15:20:04Z")

</div>

maybe we can use the new **zlib-ng** lib.

```julia
x86-64 benchmarks:
Zlib-ng is about 4x faster than zlib, and 2.1x faster than gzip for compression.*
Zlib-ng is about 2.4x faster than zlib and 1.8x faster than gzip when decompressing."*

```

- [https://github.com/zlib-ng/zlib-ng/releases/tag/2.0.0](https://github.com/zlib-ng/zlib-ng/releases/tag/2.0.0)
- via Phoronix: [Zlib-ng 2.0 Released As More Performant + Modern Zlib Fork - Phoronix](https://www.phoronix.com/scan.php?page=news_item&px=zlib-ng-2.0)

[Next page](https://discourse.julialang.org/t/slow-gzip-streaming-in-julia-but-not-in-python/56346.md?page=2)
