# .csv number of rows

**URL:** <https://discourse.julialang.org/t/csv-number-of-rows/27475>\
**Category:** Data\
**Tags:** csv\
**Created:** [August 12, 2019, 11:18pm UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475 "2019-08-12T23:18:18Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![austinbean](https://avatars.discourse-cdn.com/v4/letter/a/839c29/32.png) [@austinbean](https://discourse.julialang.org/u/austinbean)\
**Post date:** [August 12, 2019, 11:18pm UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/1 "2019-08-12T23:18:18Z")

</div>

Is there a clever way to determine the number of rows in large .csv or .tsv file without reading a whole column?

I’m currently doing:

```julia
CSV.File("big_file.csv") |> Tables.select(:first_column_name) |> DataFrame

```

which tells me what I need to know, but if there’s something which just counts the rows a little faster, I would love to know.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 12, 2019, 11:34pm UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/2 "2019-08-12T23:34:08Z")

</div>

eh… you can just find out the number of lines in the file instead…?

```julia
julia> (open("./Final.csv") |> readlines |> length) -1
1936

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [August 13, 2019, 12:32am UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/3 "2019-08-13T00:32:55Z")

</div>

Be careful with this approach, it will give wrong answer for CSV files having linebreaks in some values.

---

<div class="post-metadata">

**Author:** ![austinbean](https://avatars.discourse-cdn.com/v4/letter/a/839c29/32.png) [@austinbean](https://discourse.julialang.org/u/austinbean)\
**Post date:** [August 13, 2019, 1:14am UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/4 "2019-08-13T01:14:52Z")

</div>

thanks I figured it was easy but basic googling did not reveal it.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 13, 2019, 3:16am UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/5 "2019-08-13T03:16:36Z")

</div>

I always think that `wc -l` is much faster but of course windows doesn’t come with such things pre-built, which is a shame.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [August 13, 2019, 4:30pm UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/6 "2019-08-13T16:30:01Z")

</div>

As @aplavin mentioned, just doing readlines can be incorrect for csv files w/ quoted newline characters. Using the `readlines` function is also pretty wasteful and will gobble up a lot of memory for really large files. In Base, the `countlines` function will be much more efficient.

For a more general purpose solution for csv files that may contain quoted newline characters, this should be extremely fast/efficient:

```julia
function countcsvlines(file)
    n = 0
    for row in CSV.Rows(file; resusebuffer=true)
        n += 1
    end
    return n
end

```

---

<div class="post-metadata">

**Author:** ![mrip](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mrip](https://discourse.julialang.org/u/mrip)\
**Post date:** [September 13, 2022, 9:41pm UTC](https://discourse.julialang.org/t/csv-number-of-rows/27475/7 "2022-09-13T21:41:16Z")

</div>

UPDATE: @quinnj spells it “reSusebuffer” = true)  
I have found that “reusebuffer” = true) works much better

But other than that…  
Amazing. This is gonna help so much with pre-allocating. Loading this and a first row reader into every CSV analysis from now on!!!
