# Apply a column of anonymous functions for each column in a column subset

**URL:** <https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402>\
**Category:** Data\
**Tags:** dataframes\
**Created:** [April 12, 2022, 7:37pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402 "2022-04-12T19:37:51Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [April 12, 2022, 7:37pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/1 "2022-04-12T19:37:51Z")

</div>

Hi,

Suppose we have a DataFrame like this:

```julia
df = DataFrame(name='a':'c',x1=1:3,x2=[[1,2],[3,4],[5,6]],xfun=[x->x.-1,x->x.^2,x->x.^3])

```

I want to transform column `x1` and `x2` in `df` by apply each fun in xfun for the corresponding row of `x1` and `x2`, I could use [:x1,:xfun]=\>ByRow((x,f)-\>f(x))=\>:x1, but what if there are 20 of these columns, is there other elegant way to achieve this?

The other way i can think of is to convert the columns to a Matrix, and broadcasting a vector of anonymous functions to the first dimention of the Matrix, but i don’t know if there is a generic `apply` function to broadcast?

Thanks,  
Alex

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 12, 2022, 8:42pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/2 "2022-04-12T20:42:57Z")

</div>

```julia
julia> combine(df, vcat.(["x1", "x2"], "xfun") .=> ByRow((x,f) -> f(x)) => first)
3×2 DataFrame
 Row │ x1 x2
     │ Int64 Array…
─────┼───────────────────
   1 │ 0 [0, 1]
   2 │ 4 [9, 16]
   3 │ 27 [125, 216]

```

and instead of `["x1", "x2"]` provide an expression that generates the column names you want to include.

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [April 12, 2022, 11:37pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/3 "2022-04-12T23:37:34Z")

</div>

Thanks, it’s exactly what i want. The `transform` version works too, like this:

```julia
transform(df, vcat.(["x1", "x2"], "xfun") .=> ByRow((x,f) -> f(x)) => first)

```

is there other different between these two?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 13, 2022, 7:28am UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/4 "2022-04-13T07:28:13Z")

</div>

The differences are:

- `transform` keeps all source columns always; `combine` only keeps columns specified in transformations;
- `transform` requires output to have as many rows as input; `combine` allows any number of rows in output.

Other than that these functions interpret transformation specifications in the same way (i.e. the same engine processes both requests, but different additional constraints are added)

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [April 13, 2022, 3:57pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/5 "2022-04-13T15:57:56Z")

</div>

just a slightly different way of combining 😉 things

```julia
cols=["x1", "x2"]
combine(df, ["xfun";cols]=>ByRow((f,x...)->f.(x))=>cols)

```

but above all to ask for information on the use of the `first` function instead of a list of names / symbols of columns in output.

PS

I wonder if and when it will also be possible to write something like this

```julia
combine(df, [cols;"xfun"]=>ByRow((x...,f)->f.(x))=>cols)

# so for the given df is possible to save some typing :-)

combine(df, 2:4=>ByRow((x...,f)->f.(x))=>2:3)

```

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [April 13, 2022, 10:41pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/6 "2022-04-13T22:41:36Z")

</div>

Thanks for the clarification!

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [April 13, 2022, 10:43pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/7 "2022-04-13T22:43:46Z")

</div>

splitting to a vector of names is also quite concise.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 14, 2022, 6:51am UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/8 "2022-04-14T06:51:04Z")

</div>

> [@rocco\_sprmnt21](#):
>
> I wonder if and when it will also be possible to write something like this

Base Julia does not allow this and I do not think it will be allowed.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [April 14, 2022, 11:48am UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/9 "2022-04-14T11:48:46Z")

</div>

I take this opportunity to ask you a further question, this one more specific one relating to the mini language.  
If I understand correctly, some input forms such as columns range are not allowed in output.  
For example 2: 3 =\> fun =\> 2: 3, it doesn’t work.  
If so, what is the reason for these restrictions?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 14, 2022, 11:58am UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/10 "2022-04-14T11:58:27Z")

</div>

> [@rocco\_sprmnt21](#):
>
> For example 2: 3 =\> fun =\> 2: 3, it doesn’t work.

This could work and would mean the following:

> pass contents of columns 2 and 3 as positional arguments to function `fun` and expand the result returned by it into two columns whose names are taken as names of columns 2 and 3 from the source

The first question is if this is what you would expect. If this is what you would expect, at least for me this is a very specific case that is needed quite rarely and currently it can be expressed as `2:3 => fun => names(df, 2:3)` which is only a bit more verbose.

For single column transformations like `2 => fun => 2` in your proposed notation, which are more common, either pass `renamecols=false` as kwarg and write just `2 => fun` or write `2 => fun => identity` to retain source column name. This does not cover the case like `2 => fun => 3`, but again I think that it is quite rare.

What is your use case where you require this kind of transformations?

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [April 14, 2022, 12:43pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/11 "2022-04-14T12:43:51Z")

</div>

the simple one: the first.  
Obviously when I did the test I mixed something else.  
Thanks

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 14, 2022, 1:02pm UTC](https://discourse.julialang.org/t/apply-a-column-of-anonymous-functions-for-each-column-in-a-column-subset/79402/12 "2022-04-14T13:02:50Z")

</div>

For a reference here is an example where your original syntax could be useful:

```julia
julia> using DataFrames

julia> fun(x, y) = map((a, b) -> (a+b, a-b), x, y)
fun (generic function with 1 method)

julia> df = DataFrame(a=1:3, b=4:6)
3×2 DataFrame
 Row │ a b
     │ Int64 Int64
─────┼──────────────
   1 │ 1 4
   2 │ 2 5
   3 │ 3 6

julia> combine(df, [:a, :b] => fun => [:a, :b])
3×2 DataFrame
 Row │ a b
     │ Int64 Int64
─────┼──────────────
   1 │ 5 -3
   2 │ 7 -3
   3 │ 9 -3

```
