# Is there an equivalent of eachindex() for DataFrames?

**URL:** <https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945>\
**Category:** General Usage\
**Tags:** question, dataframes, type-stability\
**Created:** [October 19, 2022, 7:17am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945 "2022-10-19T07:17:21Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [October 19, 2022, 7:17am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/1 "2022-10-19T07:17:22Z")

</div>

Hi suppose I have an array and want append some string to entry I can iterate across it with eachindex()

```julia
singers = [["Marley", "Kiiara", "Sinead"] ["Rynn","Illenium","Nelly"]]

for k = eachindex(singers)
singers[k] = "$(singers[k]) sold out."
end 

```

Just curious if there was a similar function or way to achieve this achieve this with a DataFrame?  
If I try something like

```julia
df = DataFrame(singers, :auto) 

for (i,j) = (eachrow(singersdf),eachcol(singersdf))
       singersdf[i , j] = "$(singersdf[i , j]) today"
       end

```

Julia returns a MethodError. Thank you!

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 19, 2022, 7:52am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/2 "2022-10-19T07:52:52Z")

</div>

One option might be:

```julia
for i in 1:nrow(df), j in 1:ncol(df)
    df[i,j] = "$(df[i,j]) sold out"
end

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 19, 2022, 9:57am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/3 "2022-10-19T09:57:08Z")

</div>

`eachindex` is not supoprted for data frame. You can use what @rafael.guerra proposed. However, the main reason why it is not supported is that such indexing is very inefficient so it will have a reasonable performance only for small data frames. If you need to do such iteration use function barrier.

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [October 20, 2022, 1:09am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/4 "2022-10-20T01:09:37Z")

</div>

Thanks so much this is very helpful. Just to clarify the performance boosting methods you discussed in your Efficiency of DataFrame Row Iteration blog posts only applies if I am iterating over the rows of multiple or all columns of a DataFrame? So if I were just iterating over the rows of a single column (say each row of df.x1) then the performance boost would not apply because at that point I am just iterating over a vector?

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [October 20, 2022, 1:12am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/5 "2022-10-20T01:12:32Z")

</div>

thanks so much! this is what I was thinking about but didn’t know the correct format.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 20, 2022, 7:05am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/6 "2022-10-20T07:05:34Z")

</div>

To have high performance you need to use function barrier. Here is a most basic example:

```julia
function_barrier(vec) = ... your code iterating elements of a vector
map(function_barrier, eachcol(df))

```

The point is that in order to be efficient you must pass a column to a separate function. Then inside this function all will be fast.

The reason is that `DataFrame` object is not type stable, so for example even:

```julia
for col in eachcol(df)
    for v in col
        ... your code
    end
end

```

will be slow, because Julia does not know the element type of `col` at compilation time.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 20, 2022, 7:25am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/7 "2022-10-20T07:25:07Z")

</div>

> [@bkamins](#):
>
> will be slow, because Julia does not know the element type of `col` at compilation time.

In this case, as all data frame elements are strings, it should be faster to create a matrix of strings, iterate over each index of this matrix and edit the strings, and finally convert the result back to a DataFrame for further work?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 20, 2022, 7:39am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/8 "2022-10-20T07:39:58Z")

</div>

If all columns have the same type then what is enough is:

```julia
for col in eachcol(df)
    for v in col::Vector{String} # assuming this is the type of column
        ... your code
    end
end

```

of course converting to a `Matrix` or to `Tables.columntable` also will work in this case.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [October 20, 2022, 8:12am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/9 "2022-10-20T08:12:10Z")

</div>

I might be missing something here but I would probably write:

```julia
function f(x)
    ... your code
end

f.(eachcol(df))

```

which should get around the problem?

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 20, 2022, 8:14am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/10 "2022-10-20T08:14:21Z")

</div>

In the general case, we would need a function of both indices: `(i,j) -> f(i,j)`

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 20, 2022, 12:00pm UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/11 "2022-10-20T12:00:38Z")

</div>

So it would be `f.(enumerate(eachcol(df)))` or `f.(pairs(eachcol(df)))` depending on what kind of column index user wants.

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [October 20, 2022, 11:41pm UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/12 "2022-10-20T23:41:21Z")

</div>

Thanks for all the explanations and additional methods! Sorry I am still a little confused about the distinction between `f.(eachcol(df))`, `f.(enumerate(eachcol(df)))`, and `f.(pairs(eachcol(df)))` and how to implement the latter two.

```julia
for col in eachcol(df)
           for v in col::Vector{String} 
               println("$v today")
           end
       end

```

gives the desired result of

```julia
Marley sold out today
Kiiara sold out today
Sinead sold out today
Rynn sold out today
Illenium sold out today
Nelly sold out today

```

Likewise if I define

```julia
function f(x)
     for i = x
    println("$i today")
end

```

then both

```julia
julia> f.(eachcol(df))

```

and

```julia
map(f,eachcol(df))

```

yield

```julia
Marley sold out today
Kiiara sold out today
Sinead sold out today
Rynn sold out today
Illenium sold out today
Nelly sold out today
2-element Vector{Nothing}:
 nothing
 nothing

```

but

```julia
julia> f.(enumerate(eachcol(df)))

1 today
["Marley sold out", "Kiiara sold out", "Sinead sold out"] today
2 today
["Rynn sold out", "Illenium sold out", "Nelly sold out"] today
2-element Vector{Nothing}:
 nothing
 nothing

```

and

```julia
julia> f.(pairs(eachcol(df)))

ERROR: ArgumentError: broadcasting over dictionaries and `NamedTuple`s is reserved
Stacktrace:
 [1] broadcastable(#unused#::Base.Pairs{Symbol, AbstractVector, Vector{Symbol}, DataFrames.DataFrameColumns{DataFrame}})
   @ Base.Broadcast ./broadcast.jl:705
 [2] broadcasted(::Function, ::Base.Pairs{Symbol, AbstractVector, Vector{Symbol}, DataFrames.DataFrameColumns{DataFrame}})
   @ Base.Broadcast ./broadcast.jl:1295
 [3] top-level scope
   @ REPL[309]:1

```

Is this because with `enumerate` and `pairs` it is no longer a vector being inputted into the function? How does the function need to be modified for it to work? Sorry I’m sure I’m missing something simple here. I though maybe removing the row iteration in the function might work but

```julia
function g(x)
   println("$x today")
end

```

yields

```julia
g.(enumerate(eachcol(df)))

(1, ["Marley sold out", "Kiiara sold out", "Sinead sold out"]) today
(2, ["Rynn sold out", "Illenium sold out", "Nelly sold out"]) today
2-element Vector{Nothing}:
 nothing
 nothing

```

and

```julia
g.(pairs(eachcol(df)))

ERROR: ArgumentError: broadcasting over dictionaries and `NamedTuple`s is reserved
Stacktrace:
 [1] broadcastable(#unused#::Base.Pairs{Symbol, AbstractVector, Vector{Symbol}, DataFrames.DataFrameColumns{DataFrame}})
   @ Base.Broadcast ./broadcast.jl:705
 [2] broadcasted(::Function, ::Base.Pairs{Symbol, AbstractVector, Vector{Symbol}, DataFrames.DataFrameColumns{DataFrame}})
   @ Base.Broadcast ./broadcast.jl:1295
 [3] top-level scope
   @ REPL[314]:1

```

other modifications to the function I’ve tried also yields errors.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 21, 2022, 7:07am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/13 "2022-10-21T07:07:22Z")

</div>

Let me give you a simpler example of the difference:

```julia
julia> df = DataFrame(a=1, b=2, c=3)
1×3 DataFrame
 Row │ a b c
     │ Int64 Int64 Int64
─────┼─────────────────────
   1 │ 1 2 3

julia> collect(eachcol(df))
3-element Vector{AbstractVector}:
 [1]
 [2]
 [3]

julia> collect(enumerate(eachcol(df)))
3-element Vector{Tuple{Int64, AbstractVector}}:
 (1, [1])
 (2, [2])
 (3, [3])

julia> collect(pairs(eachcol(df)))
3-element Vector{Pair{Symbol, AbstractVector}}:
 :a => [1]
 :b => [2]
 :c => [3]

```

So, as you can see, the difference is just hat you have different objects returned. In the `eachindex` case you get column number as a first element. In the `pairs` case you get column name as a first element.

As for broadcasting not working for `pairs` - I have forgotten that `pairs` returns `AbstractDict`, so in this case you need to use `foreach` instead.

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [October 21, 2022, 10:18am UTC](https://discourse.julialang.org/t/is-there-an-equivalent-of-eachindex-for-dataframes/88945/14 "2022-10-21T10:18:13Z")

</div>

great thanks for clearing that up!
