# Handle large csv file using \`enumerate(CSV.File())\` or \`CSV.read()\`?

**URL:** <https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309>\
**Category:** New to Julia\
**Created:** [April 19, 2019, 4:25pm UTC](https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309 "2019-04-19T16:25:54Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![bsnyh](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@bsnyh](https://discourse.julialang.org/u/bsnyh)\
**Post date:** [April 19, 2019, 4:25pm UTC](https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309/1 "2019-04-19T16:25:55Z")

</div>

I need to read a large .CSV file and parse each line to the attributes of a self-defined type. Before handling the CSV file, I was thinking to create a vector of a fixed length, which is the number of rows of the CSV file. Based on the past experiences, I thought that having the vector length fixed would be more efficient, in terms of vector element assigning. So, is there anyway to get to know the number of rows of the .CSV file without reading the whole file? Probably not, right? How do you manage to design your code to avoid this?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [April 19, 2019, 5:07pm UTC](https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309/2 "2019-04-19T17:07:30Z")

</div>

I would just `push!` rows as a known type, probably a `NamedTuple` or a `struct`, to a `Vector`. The implementation of `push!` is very efficient (it preallocates) and is almost surely better than reading the file twice (which is the only way to establish the number of rows).

---

<div class="post-metadata">

**Author:** ![bsnyh](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@bsnyh](https://discourse.julialang.org/u/bsnyh)\
**Post date:** [April 19, 2019, 5:33pm UTC](https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309/3 "2019-04-19T17:33:27Z")

</div>

Thank you. Or on the other hand, if I look at the input files, and take a look at the number of rows there.

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [April 21, 2019, 2:23pm UTC](https://discourse.julialang.org/t/handle-large-csv-file-using-enumerate-csv-file-or-csv-read/23309/4 "2019-04-21T14:23:03Z")

</div>

> [@bsnyh](#):
>
> So, is there anyway to get to know the number of rows of the .CSV file without reading the whole file? Probably not, right? How do you manage to design your code to avoid this?

The CSV format relies on the newline character (and carriage return on Windows) to separate each record. Further, every record may have a different length. So, there is no way to find how many rows in the file without examining every single character in the file.

If you have a very large file then you can parallelize the read/count across multiple workers to speed that up.

@Tamas_Papp had a good suggestion for using `push!` so you can avoid this work.
