# Transforming dataframe columns selected via regex while keeping the other columns

**URL:** <https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702>\
**Category:** Data\
**Tags:** regex, dataframes\
**Created:** [May 24, 2021, 3:45am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702 "2021-05-24T03:45:35Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon71627001](https://avatars.discourse-cdn.com/v4/letter/a/a698b9/32.png) [@anon71627001](https://discourse.julialang.org/u/anon71627001)\
**Post date:** [May 24, 2021, 3:45am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702/1 "2021-05-24T03:45:35Z")

</div>

Let’s say I have `df = DataFrame(ax = 1:3, bx = 4:6, cy = 7:9, dy = 10:12)` and want to double the values of each column that has an “x” in its header, while leaving the columns that don’t have an “x” in their header unchanged. When I try to do that with

`transform(df, r"x" => ByRow(x -> 2x); renamecols = false)`

I get `MethodError: no method matching (::Main.workspace4.var"#1#2")(::Int64, ::Int64)`. It looks like it’s trying to use them each as inputs to a single bivariate function, because the following code runs

`transform(df, r"x" => ByRow(+))`

I know I could select columns with regex and then apply the transformation to all of them, but that would leave the other columns out, which I don’t want.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [May 24, 2021, 7:37am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702/2 "2021-05-24T07:37:00Z")

</div>

you should know (which I don’t know) what the expression r “x” produces in this context (a tuple perhaps?) , in order to handle it properly with the function x-\> 2x.  
While waiting for an explanation of what is going owhen one use a regex expression to select columns, you can try using such a workaround …

```julia
transform(df, [x for x in names(df) if contains(x,"x")].=> ByRow(x -> 2x);renamecols=false)

```

pay attention to the “.” in “. =\>”

anhoter way, using the named tuple’s properties

```julia
transform(df, AsTable(r"x")=> (nt->(;zip(keys(nt),2 .* values(nt))...))=>AsTable)

```

or using splatting operator

```julia
transform(df, r"x".=> ByRow((x...) -> 2 .*x)=>filter(x->contains(x,"x"),names(df)))

```

or this way

```julia
transform(df, r"x".=> ByRow((x...) -> 2 .*x)=>names(select(df,r"x")))

```

but unfortunately, the following doesn’t work

```julia
transform(df, r"x".=> ByRow((x...) -> 2 .*x)=>r"x")

```

_MethodError: no method matching getindex(::DataFrames.Index, ::Pair{Regex, Pair{ByRow{var"#181#182"}, Regex}})_

at least not in this naive form. Perhaps using the regex capabilities appropriately, the result can be achieved.

---

<div class="post-metadata">

**Author:** ![qsong](https://avatars.discourse-cdn.com/v4/letter/q/d07c76/32.png) [@qsong](https://discourse.julialang.org/u/qsong)\
**Post date:** [May 24, 2021, 8:32am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702/3 "2021-05-24T08:32:26Z")

</div>

Or

```julia
transform(df, names(df, r"x") .=> (x -> 2x), renamecols=false)

```

---

<div class="post-metadata">

**Author:** ![anon71627001](https://avatars.discourse-cdn.com/v4/letter/a/a698b9/32.png) [@anon71627001](https://discourse.julialang.org/u/anon71627001)\
**Post date:** [May 24, 2021, 9:52am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702/5 "2021-05-24T09:52:05Z")

</div>

Thanks for the solution. Is there a way to do it using `ByRow` , in case the applied function operates on each element of the column instead of the entire column?

---

<div class="post-metadata">

**Author:** ![qsong](https://avatars.discourse-cdn.com/v4/letter/q/d07c76/32.png) [@qsong](https://discourse.julialang.org/u/qsong)\
**Post date:** [May 24, 2021, 10:03am UTC](https://discourse.julialang.org/t/transforming-dataframe-columns-selected-via-regex-while-keeping-the-other-columns/61702/6 "2021-05-24T10:03:40Z")

</div>

`ByRow` is the other way to do the “broadcasting” for dataframes like you said. Take the addition for example. We will get the same result by either of the following

```julia
transform(df, names(df, r"x") .=> (x -> 100 .+ x), renamecols=false)
transform(df, names(df, r"x") .=> ByRow(x -> 100 + x), renamecols=false)

```
