# Reading large csv file

**URL:** <https://discourse.julialang.org/t/reading-large-csv-file/67258>\
**Category:** Data\
**Tags:** performance, csv\
**Created:** [August 29, 2021, 1:25am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258 "2021-08-29T01:25:04Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 29, 2021, 1:25am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/1 "2021-08-29T01:25:05Z")

</div>

I checked some answers and all point to CSV.jl, but it is slow(i guess this is “time to first plot” issue), it uses a lot of memory, it produces some strange type of array as output… Do I have another choice?

---

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 29, 2021, 1:27am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/2 "2021-08-29T01:27:15Z")

</div>

some times CSV.jl kills julia (out of memory issue)

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 29, 2021, 1:27am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/3 "2021-08-29T01:27:44Z")

</div>

Don’t worry about the types of output. They will work just like any other julia array.

Why is the time to first plot issue a problem? This might mean you are calling a julia script from the command line over and over again, which is not recommended.

How big is the CSV file? Is it bigger than your RAM?

---

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 29, 2021, 1:29am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/4 "2021-08-29T01:29:26Z")

</div>

the final result fits in memory

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 29, 2021, 1:30am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/5 "2021-08-29T01:30:34Z")

</div>

hmmm… I’m not sure, then. Maybe you need to specify the types via a keyword argument more strictly? What version of CSV.jl are you on?

---

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 29, 2021, 1:31am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/6 "2021-08-29T01:31:27Z")

</div>

I am already struggling with pooled array and categorical array 🙃 now there is something like chain array ?! 😬

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 29, 2021, 1:35am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/7 "2021-08-29T01:35:55Z")

</div>

Chain array is a way to read in the data faster. I really wouldn’t worry about it, these are optimizations that should be hidden from the user to maximize performance.

Pooled array is the same way. It’s only there for performance (saving memory).

There seem to be a few keyword arguments that can reduce the memory footprint in the [documentation](https://csv.juliadata.org/stable/#Getting-Started). Maybe those will help.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 29, 2021, 1:37am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/8 "2021-08-29T01:37:20Z")

</div>

any example files?

---

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 29, 2021, 1:46am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/9 "2021-08-29T01:46:08Z")

</div>

is there a package to read and write stata file (stata 16). maybe it is better than csv?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 29, 2021, 1:47am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/10 "2021-08-29T01:47:46Z")

</div>

Maybe [this](https://github.com/junyuan-chen/ReadStatTables.jl)? It uses the same C library as R’s `haven` does iirc.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [August 29, 2021, 7:08am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/11 "2021-08-29T07:08:18Z")

</div>

Do you have a CSV file which is read in successfully and fast in R, Python or similar but that CSV.jl reads in very slowly or even not at all? If so I’m sure that would be considered a bug in CSV.jl so would be great to reproduce.

Time to first plot is unlikely to be an issue here if the file is very large, ie the parsing actually takes nonnegligible time.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 29, 2021, 8:53am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/12 "2021-08-29T08:53:10Z")

</div>

Additionally CSV.jl 1.0 release will be soon announced, where @quinnj wants to resolve “time to first plot” as much as possible. As @nilshg commented - if you run into issues please open a reproducible issue in CSV.jl and it will be handled.

Finally, in order to use “standard” types use `pool=false` and `threaded=false`, as this will turn-off most of the optimizations that cause you issues (note though that `pool=false` will disable pooling of string columns which will increase the memory footprint of the object you read in).

---

<div class="post-metadata">

**Author:** ![xinchin](https://avatars.discourse-cdn.com/v4/letter/x/54ee81/32.png) [@xinchin](https://discourse.julialang.org/u/xinchin)\
**Post date:** [August 30, 2021, 12:10am UTC](https://discourse.julialang.org/t/reading-large-csv-file/67258/13 "2021-08-30T00:10:20Z")

</div>

threaded = false helped (for me it is much faster than not setting it), so i should set it false when I am reading large files?
