# Reading huge csv files

**URL:** <https://discourse.julialang.org/t/reading-huge-csv-files/15924>\
**Category:** Data\
**Created:** [October 5, 2018, 1:27pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924 "2018-10-05T13:27:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![jayce\_ram](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jayce_ram/32/3977_2.png) [@jayce\_ram](https://discourse.julialang.org/u/jayce_ram)\
**Post date:** [October 5, 2018, 1:27pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/1 "2018-10-05T13:27:55Z")

</div>

I would like to read a huge csv file around 9G and apply filters to the rows in order to build another dataframe carrying only the values I want to. I have installed the pack DataBase and I have just tried the simple command :

```julia
using JuliaDB 
flights = loadtable("Document.csv")

```

and I got the following error:

```julia
OutOfMemoryError()

```

How I can deal with request of this type on Julia? Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 5, 2018, 1:40pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/2 "2018-10-05T13:40:08Z")

</div>

> [@jayce\_ram](#):
>
> How I can deal with request of this type on Julia? Thanks in advance.

Unfortunately, handling large data in Julia is difficult at the moment. I would recommend opening a `CSV.File`, iterating through it, and only keeping what you need. Look at the docstrings in `CSV.jl`.

---

<div class="post-metadata">

**Author:** ![jayce\_ram](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jayce_ram/32/3977_2.png) [@jayce\_ram](https://discourse.julialang.org/u/jayce_ram)\
**Post date:** [October 5, 2018, 1:47pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/3 "2018-10-05T13:47:59Z")

</div>

Thanks for your answer. I was wondering about your suggestion: I cannot use `CSV.read` because this function cannot neither deal with this huge file csv. Maybe I didn’t understand well your suggestion. Could you clarify it to me ? Thanks.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 5, 2018, 2:00pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/4 "2018-10-05T14:00:06Z")

</div>

This is the docstring I suggested you read:

[https://github.com/JuliaData/CSV.jl/blob/5c0b6a8b6f2eb9f30c36c8dc3754137a1deed68f/src/CSV.jl#L49](https://github.com/JuliaData/CSV.jl/blob/5c0b6a8b6f2eb9f30c36c8dc3754137a1deed68f/src/CSV.jl#L49)

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 10, 2018, 2:33am UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/5 "2018-10-10T02:33:58Z")

</div>

As @Tamas_Papp mentioned, using `CSV.File(file; kw...)` will return a `CSV.File` object which doesn’t load the entire dataset into memory. You could then “build up” a table by iterating over the rows and filtering as you’d like, something like:

```julia
function buildtable(filter::Function, file)
    f = CSV.File(file)
    # create a NamedTuple of Vectors to push! to
    table = (colA=Int[], colB=Float64[], colC=String[])
    for row in f
        if filter(row)
            push!(table.colA, row.colA)
            push!(table.colB, row.colB)
            push!(table.colC, row.colC)
        end
    end
    return table
end

```

---

<div class="post-metadata">

**Author:** ![Wikunia](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wikunia/32/2180_2.png) [@Wikunia](https://discourse.julialang.org/u/Wikunia)\
**Post date:** [January 19, 2019, 10:52pm UTC](https://discourse.julialang.org/t/reading-huge-csv-files/15924/6 "2019-01-19T22:52:29Z")

</div>

This doesn’t seem to be true:

```julia
function readtest()
    i = 1
    for row in CSV.File("abc.csv"; allowmissing=:none)   
        i += 1
        if i > 5
            break
        end
    end  
end
readtest()

```

where abc.csv is ~10GB doesn’t work. If I use

```julia
CSV.File("abc.csv"; allowmissing=:none, limit=5)

```

it works.  
It seems it’s trying to read the whole dataset in the first case or am I wrong?
