# Rewriting dplyr code which uses a function of columns in Julia -style using DataFrames.jl

**URL:** <https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [March 25, 2021, 3:35pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939 "2021-03-25T15:35:18Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Adam\_Fleischhacker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adam_fleischhacker/32/24468_2.png) [@Adam\_Fleischhacker](https://discourse.julialang.org/u/Adam_Fleischhacker)\
**Post date:** [March 25, 2021, 3:35pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/1 "2021-03-25T15:35:18Z")

</div>

I am migrating my brain towards thinking in Julia-esque style. Yet, I am not sure how to approach this problem where I want to create a third column of a data frame that is a function of the other two.  
I prefer intuitive/readable code and am wondering if I should add another method to my function, use broadcasting, use byRow, or what would be recommended. The anonymous function syntax that I see seems overly verbose, but I am open to that if that is the way.

```julia
## reproducible Example
using DataFrames

## define function of x and y
function foo(x::Float64,y::Float64)
    2*x*y
end

## create data frame with three columns
DataFrame(colX = rand(3),
        colY = rand(3),
        colZ = foo.(:colX,:colY))

# ERROR: LoadError: MethodError: no method matching foo(::Symbol, ::Symbol)

```

How would you rewrite the above in a way that works and is highly readable. My point of comparison is this R code:

```julia
library(dplyr)
foo = function(x,y) {
  2 * x * y
}

tibble(x = runif(3),
       y = runif(3),
       z = foo(x,y))

```

---

<div class="post-metadata">

**Author:** ![chris-b1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chris-b1/32/14165_2.png) [@chris-b1](https://discourse.julialang.org/u/chris-b1)\
**Post date:** [March 25, 2021, 3:55pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/2 "2021-03-25T15:55:04Z")

</div>

With DataFrames.jl, you can’t reference other columns by name during construction, and in general don’t have the non-standard evaluation that dpylr uses. Two options for writing would be

```julia
# 1. make the vectors first
y = rand(3)
x = rand(3)
df = DataFrame(x=x, y=y, z=foo.(x, y))

# 2. add the column after the fact
df = DataFrame(x=rand(3), y=rand(3))
df.z = foo.(df.x, df.y)

```

---

<div class="post-metadata">

**Author:** ![ElOceanografo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eloceanografo/32/624_2.png) [@ElOceanografo](https://discourse.julialang.org/u/ElOceanografo)\
**Post date:** [March 25, 2021, 4:45pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/3 "2021-03-25T16:45:58Z")

</div>

I find that [DataFramesMeta](https://github.com/JuliaData/DataFramesMeta.jl) gives me the closest experience to `dplyr`:

```julia
using DataFrames, DataFramesMeta
foo(x, y) = 2 * x * y
df = DataFrame(x=randn(10), y=randn(10))
# using the @transform macro 
@transform(df, z = foo.(:x, :y))
# ...or chaining operations using @linq
@linq df |> transform(z = foo.(:x, :y))

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 25, 2021, 5:17pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/4 "2021-03-25T17:17:30Z")

</div>

I think DataFramesMeta + Chain.jl gives the most dply-like experiience

```julia
julia> df = DataFrame(x=randn(10), y=randn(10));

julia> @chain df begin 
           @transform(y = foo.(:x, :y))
           @transform(g = rand(0:1, length(:y)))
           groupby(:g)
           @combine(y_mean = mean(:y), x_mean = mean(:x))
       end
2×3 DataFrame
 Row │ g y_mean x_mean    
     │ Int64 Float64 Float64   
─────┼────────────────────────────
   1 │ 0 0.967425 -0.513779
   2 │ 1 1.68032 -0.906092

```

---

<div class="post-metadata">

**Author:** ![Adam\_Fleischhacker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adam_fleischhacker/32/24468_2.png) [@Adam\_Fleischhacker](https://discourse.julialang.org/u/Adam_Fleischhacker)\
**Post date:** [March 25, 2021, 5:46pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/5 "2021-03-25T17:46:40Z")

</div>

I like this paradigm the best. Very readable and melds well with my brain.

---

<div class="post-metadata">

**Author:** ![Adam\_Fleischhacker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adam_fleischhacker/32/24468_2.png) [@Adam\_Fleischhacker](https://discourse.julialang.org/u/Adam_Fleischhacker)\
**Post date:** [March 25, 2021, 5:49pm UTC](https://discourse.julialang.org/t/rewriting-dplyr-code-which-uses-a-function-of-columns-in-julia-style-using-dataframes-jl/57939/6 "2021-03-25T17:49:14Z")

</div>

I do like #2 alot, but really like the chaining workflow of dplyr as presented below. Thx.
