# Is it possible to iterate over a very large CSV in Windows?

**URL:** <https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827>\
**Category:** New to Julia\
**Created:** [August 22, 2019, 2:23am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827 "2019-08-22T02:23:40Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![gabomgp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gabomgp/32/1918_2.png) [@gabomgp](https://discourse.julialang.org/u/gabomgp)\
**Post date:** [August 22, 2019, 2:23am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/1 "2019-08-22T02:23:40Z")

</div>

I’m using CSV.jl to iterate over a very large CSV, but in Windows it’s throwing the error:

```
Error: could not create file mapping: The operation comleted successfully

```

The same error that this issue describe:

```
https://github.com/JuliaData/CSV.jl/issues/424

```

The error appears with CSV.File and CSV.Rows, so i’m thinking that in Windows is not possible to iterate over very large CSV’s. Do exists another option to do that task?

I don’t need type inference, i can specify the name/type of every column, and the separator if necesary.

---

<div class="post-metadata">

**Author:** ![gabomgp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gabomgp/32/1918_2.png) [@gabomgp](https://discourse.julialang.org/u/gabomgp)\
**Post date:** [August 22, 2019, 2:29am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/2 "2019-08-22T02:29:41Z")

</div>

My current code is:

```julia
using Pkg

Pkg.add("CodecZlib")
Pkg.add("CSV")
Pkg.add("Queryverse")

using Query, CSV, CodecZlib, Dates

function mylength(iter)
    n=0
        for i in iter  
            n+=1
        end
    return n
end

function test(src:: String)
    open(GzipDecompressorStream, src) do stream
        table = CSV.Rows(stream; 
            header=[:day, :glnprovider, :glnretailerlocation, :gtin, :inventory, :cost, :sales, :price],
            dateformat="yyyy-mm-dd",
            types=[Date, UInt64, UInt64, UInt64, Float32, Float32, Float32, Float32],
            strict=true)
            return table |> mylength
    end
end

test("D:\\Data\\03012019_03312019_17440.csv.gz")

```

---

<div class="post-metadata">

**Author:** ![bernhard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bernhard/32/2619_2.png) [@bernhard](https://discourse.julialang.org/u/bernhard)\
**Post date:** [August 22, 2019, 11:25am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/3 "2019-08-22T11:25:18Z")

</div>

Can you post a complete example?

---

<div class="post-metadata">

**Author:** ![gabomgp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gabomgp/32/1918_2.png) [@gabomgp](https://discourse.julialang.org/u/gabomgp)\
**Post date:** [August 22, 2019, 2:06pm UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/4 "2019-08-22T14:06:35Z")

</div>

I changed the question with the complete code that fails. I can’t share the data, but is a CSV with size:

- Compressed: 2.4 GB
- Uncompressed: Almost 50 GB

Can use another library for this task in Julia?

---

<div class="post-metadata">

**Author:** ![haberdashPI](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/haberdashpi/32/26337_2.png) [@haberdashPI](https://discourse.julialang.org/u/haberdashPI)\
**Post date:** [August 23, 2019, 1:00am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/5 "2019-08-23T01:00:01Z")

</div>

You could try [CSVFiles](https://github.com/queryverse/CSVFiles.jl) and see if it helps.

---

<div class="post-metadata">

**Author:** ![greg\_plowman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/greg_plowman/32/8100_2.png) [@greg\_plowman](https://discourse.julialang.org/u/greg_plowman)\
**Post date:** [August 23, 2019, 3:08am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/6 "2019-08-23T03:08:53Z")

</div>

Iterating with `CSV.Rows` should have a low memory footprint and handle large files.

However if you have are having issues (limited memory?) then you could try streaming using basic primitives. In your case, it seems your CSV file has known types and is not complex to parse.

So perhaps you could try something like:

```julia
using Dates
const types = [Date, UInt64, UInt64, UInt64, Float32, Float32, Float32, Float32]
const testfile = "testfile.csv"

function testwrite(filename::AbstractString)
    open(filename, "w") do io
        println(io, "2019-08-23, 5, 123, 17, 13.5, 1200.34, 1500.80, 22.30")
        println(io, "2019-08-22, 4, 122, 16, 12.5, 1100.34, 1400.80, 21.30")
    end
end

function testread(filename::AbstractString)
    nfields = length(types)
    for line in eachline(filename)
        strings = split(line, ",")
        fields = ntuple(i -> parse(types[i], strings[i]), nfields)
        println(fields)
        # do stuff with fields
    end
end

testwrite(testfile)
testread(testfile)

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 23, 2019, 11:03am UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/7 "2019-08-23T11:03:29Z")

</div>

> [@gabomgp](#):
>
> CSV.Rows(stream; header=[:day, :glnprovider, :glnretailerlocation, :gtin, :inventory, :cost, :sales, :price], dateformat=“yyyy-mm-dd”, types=[Date, UInt64, UInt64, UInt64, Float32, Float32, Float32, Float32], strict=true)

See this [Iterating over chunks efficiently · Issue #482 · JuliaData/CSV.jl · GitHub](https://github.com/JuliaData/CSV.jl/issues/482)

Perhaps setting `reusebuffer=true` will be helpful

---

<div class="post-metadata">

**Author:** ![gabomgp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gabomgp/32/1918_2.png) [@gabomgp](https://discourse.julialang.org/u/gabomgp)\
**Post date:** [August 23, 2019, 2:38pm UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/8 "2019-08-23T14:38:06Z")

</div>

The option reusebuffer=true doesn’t changed the result.

---

<div class="post-metadata">

**Author:** ![gabomgp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gabomgp/32/1918_2.png) [@gabomgp](https://discourse.julialang.org/u/gabomgp)\
**Post date:** [August 23, 2019, 2:44pm UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/9 "2019-08-23T14:44:10Z")

</div>

Thanks. I can try this, but probably i wil use another soluction (the idea is to teach the same tool for this type of task to others, and to be very easy is a possitive point)

I hoped this task was as easy as use a library. I’m thinking in Julia for use as ETL tool really, because it was easy to read, process and write data in streamming with higth performance. I suppose the ecosystem is not ready yet.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 23, 2019, 2:56pm UTC](https://discourse.julialang.org/t/is-it-possible-to-iterate-over-a-very-large-csv-in-windows/27827/10 "2019-08-23T14:56:16Z")

</div>

Yeah. JuliaDB is not ready for general IMHO. I think I want try my hands at a Julia disk.frame-like or dask-like at some point.
