# Piping DataFrame rows

**URL:** <https://discourse.julialang.org/t/piping-dataframe-rows/49576>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [November 4, 2020, 12:01pm UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576 "2020-11-04T12:01:52Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [November 4, 2020, 12:01pm UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/1 "2020-11-04T12:01:52Z")

</div>

`@pipe eachrow( my_DataFrame ) .|> _`

This gives an error:  
“Objects of type DataFrameRow{DataFrame, DataFrames.Index} are not callable”

It would be very useful to Pipe the rows of a DataFrame. Is there a way to avoid the error?

---

<div class="post-metadata">

**Author:** ![fbanning](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fbanning/32/14972_2.png) [@fbanning](https://discourse.julialang.org/u/fbanning)\
**Post date:** [November 4, 2020, 1:11pm UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/2 "2020-11-04T13:11:47Z")

</div>

Why not do

```julia
using DataFrames

df = DataFrame(a = rand(100))

for row in eachrow(df)
    row.a * 2 # or do whatever you need to do
end

```

You can also do this via piping:

```julia
@pipe eachrow(df) .|> _.a * 2

```

Of course piping is faster and needs less allocations:

```julia
julia> @benchmark for row in eachrow(df)
       row.a * 2
       end
BenchmarkTools.Trial:
  memory estimate: 20.34 KiB
  allocs estimate: 501
  --------------
  minimum time: 16.399 μs (0.00% GC)
  median time: 17.100 μs (0.00% GC)
  mean time: 19.240 μs (5.33% GC)
  maximum time: 2.709 ms (99.02% GC)
  --------------
  samples: 10000
  evals/sample: 1

julia> @benchmark @pipe eachrow(df) .|> _.a * 2
BenchmarkTools.Trial:
  memory estimate: 4.34 KiB
  allocs estimate: 210
  --------------
  minimum time: 10.399 μs (0.00% GC)
  median time: 11.600 μs (0.00% GC)
  mean time: 12.083 μs (0.96% GC)
  maximum time: 1.176 ms (98.44% GC)
  --------------
  samples: 10000
  evals/sample: 1

```

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [November 5, 2020, 3:30am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/3 "2020-11-05T03:30:22Z")

</div>

Thanks Frederik.  
That’s helpful.  
I’d like to create many new columns using row level logic, and have new columns depend on other new columns. The shortest way I’ve found is to convert the static DataFrameRow object to a dynamic Dictionary and then to a DotMap for nicer notation. This conversion takes four steps:  
DataFrameRow → NamedTuple → Dictionary → DotMap.  
Is there a better approach?

```julia
using DataFrames, Pipe, DotMaps, NamedTupleTools

df = DataFrame(a = rand(10))

@pipe eachrow(df) .|> begin  
    r = DotMap(convert(Dict,NamedTuple(_)))    
    r.New = r.a * 2
    r.New2 = r.New + 10
    r
end |> DataFrame

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [November 5, 2020, 4:20am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/4 "2020-11-05T04:20:29Z")

</div>

The logic you have don’t require row by row

```julia
df.New = 2*df.a
df.New2 = df.New .+ 10
df

```

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [November 5, 2020, 4:27am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/5 "2020-11-05T04:27:15Z")

</div>

I’m sure your problem is more complex than the MWE, but I will at least partially second what @xiaodai said. It doesn’t quite look like you need row-by-row logic because you never mentioned anything about referring to other rows – only to other columns. I would use row-by-row if I needed to keep track of the elements from the previous row, for example. If you just need to refer to other columns, it seems like you could do it “normally” and it should work fine and relatively fast.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [November 5, 2020, 4:37am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/6 "2020-11-05T04:37:01Z")

</div>

> [@tbeason](#):
>
> I’m sure your problem is more complex than the MWE

I was thinking the same thing too.

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [November 5, 2020, 4:38am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/7 "2020-11-05T04:38:05Z")

</div>

I could have 10 or more calculated columns to add. Some might require more complex operations. Its nice to put them in a block and forget about vectorising over other dimensions (columns) within the block

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [November 5, 2020, 5:06am UTC](https://discourse.julialang.org/t/piping-dataframe-rows/49576/8 "2020-11-05T05:06:18Z")

</div>

> [@Lincoln\_Hannah](#):
>
> forget about vectorising over other dimensions (columns) within the block

that might be slow for large datasets, but for smaller datasets it should be fine.
