# Reading compressed csv gzopen cuts file short

**URL:** <https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885>\
**Category:** General Usage\
**Tags:** dataframes, io, data-compression\
**Created:** [July 1, 2026, 7:57am UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885 "2026-07-01T07:57:30Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [July 1, 2026, 7:57am UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885/1 "2026-07-01T07:57:30Z")

</div>

I seem to have stumbled into some kind of bug. I am working with large compressed CSV files. I’m loading the data with:  
`mydata = gzopen(x->CSV.read(x, DataFrame), mydatafile)`

Then I just get this warning:

```julia-auto
┌ Warning: thread = 1 warning: only found 82 / 88 columns around data row: 545401. Filling remaining columns with `missing`
└ @ CSV ~/.julia/packages/CSV/LiiJM/src/file.jl:592
545401×88 DataFrame

```

except the file has 719913 rows. I tried running the command again from the command line, then it first gave me 534731 rows, then afterwards it would give me 0 rows. It seems to me somewhere in the pipeline where CSV is trying to read from the gzip stream it’s failing when it should only wait some more, something like that.

Julia 1.12.5

```julia-auto
  [336ed68f] CSV v0.10.16
  [a93c6f00] DataFrames v1.8.2
  [92fee26a] GZip v0.7.1

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [July 1, 2026, 8:28am UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885/2 "2026-07-01T08:28:07Z")

</div>

Are you able to share the data? Does CSV read everything correctly when you unzip the data first?

---

<div class="post-metadata">

**Author:** ![xor0110](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xor0110/32/7926_2.png) [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Post date:** [July 1, 2026, 1:26pm UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885/3 "2026-07-01T13:26:49Z")

</div>

If I do `gzopen(read, ...)` and pass that to `CSV.read`, then it’s fine.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [July 1, 2026, 3:08pm UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885/4 "2026-07-01T15:08:45Z")

</div>

Does it work if you use CodecZlib.jl instead of GZip.jl ?

The code would be:

```julia-auto
mydata = open(x->CSV.read(x, DataFrame), GzipDecompressorStream)

```

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [July 1, 2026, 5:31pm UTC](https://discourse.julialang.org/t/reading-compressed-csv-gzopen-cuts-file-short/137885/5 "2026-07-01T17:31:48Z")

</div>

This is looking like [Properly crash/report error when writing to full filesystem · Issue #44535 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/44535)

Basically CSV.jl decompresses a compressed CSV file into a temporary file before processing it. If there is not enough space on your temporary file-system the file will silently get truncated.
