# Performance of eachrow(::DataFrame)

**URL:** <https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165>\
**Category:** Data\
**Created:** [August 24, 2023, 5:59pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165 "2023-08-24T17:59:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [August 24, 2023, 5:59pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165/1 "2023-08-24T17:59:39Z")

</div>

DataFrames are stored with each column contiguous in memory so that implementation iterating over `eachrow` will be slow. Instead search through `df.patient_id`.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [August 24, 2023, 6:34pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165/2 "2023-08-24T18:34:12Z")

</div>

Your predicate there is true — DataFrames are column contiguous — but your conclusion doesn’t necessarily follow. It’d only be the case if you _accessed_ every element of the row. But iterating over `eachrow` gives you a lazy reference to the row and it’ll only lookup/access the value you ask for. Edit: It’s unfortunately isn’t as fast, but the trouble is a type instability, not the access pattern.

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [August 24, 2023, 6:49pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165/3 "2023-08-24T18:49:40Z")

</div>

I believe you but I’m not sure how to prove it.

```julia
using DataFrames
using BenchmarkTools
f(df, x) = findfirst(r -> r.var"1" == x, eachrow(df))
g(df, x) = findfirst(==(x), df.var"1")

@benchmark f(d, 500_000) setup = (d = DataFrame(Dict(Symbol(i) => 1:10^6 for i in 1:100)))
@benchmark g(d, 500_000) setup = (d = DataFrame(Dict(Symbol(i) => 1:10^6 for i in 1:100)))

julia> @benchmark f(d, 500_000) setup = (d = DataFrame(Dict(Symbol(i) => 1:10^6 for i in 1:100)))
BenchmarkTools.Trial: 31 samples with 1 evaluation.
 Range (min … max): 90.518 ms … 113.115 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 108.633 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 106.752 ms ± 5.340 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

                                         ▃ ▃ ▃█▃ █▃ ▃  
  ▇▁▁▁▁▁▇▁▁▁▁▁▁▁▇▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▇▁▁▁▁▇▇▁▇█▇█▁▁▁▁▇▇███▇██▁▇▁▁▁▇█ ▁
  90.5 ms Histogram: frequency by time 113 ms <

 Memory estimate: 38.13 MiB, allocs estimate: 2498981.

julia> @benchmark g(d, 500_000) setup = (d = DataFrame(Dict(Symbol(i) => 1:10^6 for i in 1:100)))
BenchmarkTools.Trial: 87 samples with 1 evaluation.
 Range (min … max): 502.719 μs … 727.920 μs ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 528.437 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 538.021 μs ± 36.613 μs ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▆▄ ▂▄▄█ ▄ ▄ ▂                                           
  ██▆████▁▆█▆▄█▁█▆▄▄▆█▄█▆▆██▆▁▄▁▁▄▁▁▁▄▁▁▁▁▁▁▁▁▁▁▄▁▁▁▁▁▁▁▄▄▁▁▁▁▆ ▁
  503 μs Histogram: frequency by time 640 μs <

 Memory estimate: 32 bytes, allocs estimate: 2.

```

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [August 24, 2023, 7:16pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165/4 "2023-08-24T19:16:08Z")

</div>

The trouble is that DataFrames’ columns aren’t typed. Accessing `df.col1` up-front pays that penalty once whereas accessing `r.col1 for r in eachrow(df)` pays it on every single iteration.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 24, 2023, 7:55pm UTC](https://discourse.julialang.org/t/performance-of-eachrow-dataframe/103165/5 "2023-08-24T19:55:02Z")

</div>

Yes `eachrow` is designed to be convenient not fast (because of the type-instability).  
There are many use cases, when `eachrow` is fast enough and the 5x overhead reported by @jar1 above is negligible from the user’s perspective.

If someone wants a type stable iterator (that is faster) then `Tables.namedtupleiterator` can be used (but in this case it is slower than `eachrow` for some reason - I have not investigated it in detail why).
