# \[ANN\] DataFrameMacros v0.2

**URL:** <https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651>\
**Category:** Package Announcements\
**Tags:** dataframes\
**Created:** [December 26, 2021, 4:16pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651 "2021-12-26T16:16:48Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 26, 2021, 4:16pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/1 "2021-12-26T16:16:48Z")

</div>

With DataFrames 1.3, it’s now possible to mutate subsets of DataFrames. The normal approach looks like this:

```julia
df = DataFrame(x = 1:5, y = [2, 4, 6, missing, 10])
sdf = subset(df, :y => ByRow(ismissing), view = true)
transform!(sdf, :x => ByRow(x -> 2x) => :y)
df

```

Because this is a multi-step approach it doesn’t work well with piping / Chain.jl, so it’s natural to introduce some kind of special macro solution for [DataFrameMacros.jl](https://github.com/jkrumbiegel/DataFrameMacros.jl/) v0.2. I’ve iterated back and forth on this and now decided on an optional `@subset` argument to `@transform!` and `@select!`. (The `@subset` expression by itself could not be executed without a DataFrame argument, so this only works within `@transform!` and `@select!`)

## Example

```julia
julia> df = DataFrame(x = 1:5, y = [2, 4, 6, missing, 10])
5×2 DataFrame
 Row │ x y       
     │ Int64 Int64?  
─────┼────────────────
   1 │ 1 2
   2 │ 2 4
   3 │ 3 6
   4 │ 4 missing 
   5 │ 5 10

julia> @transform!(df, @subset(ismissing(:y)), :y = 2 * :x)
5×2 DataFrame
 Row │ x y      
     │ Int64 Int64? 
─────┼───────────────
   1 │ 1 2
   2 │ 2 4
   3 │ 3 6
   4 │ 4 8
   5 │ 5 10

julia> @transform!(df, @subset(:x >= 3), :z = :y + :x)
5×3 DataFrame
 Row │ x y z       
     │ Int64 Int64? Int64?  
─────┼────────────────────────
   1 │ 1 2 missing 
   2 │ 2 4 missing 
   3 │ 3 6 9
   4 │ 4 8 12
   5 │ 5 10 15

# the flag macros like `@c` for column-wise mode also work as usual

julia> @transform!(df, @subset(@c :x .< sum(:x) / length(:x)), :y = 0)
5×3 DataFrame
 Row │ x y z       
     │ Int64 Int64? Int64?  
─────┼────────────────────────
   1 │ 1 0 missing 
   2 │ 2 0 missing 
   3 │ 3 6 9
   4 │ 4 8 12
   5 │ 5 10 15

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 26, 2021, 5:10pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/2 "2021-12-26T17:10:46Z")

</div>

Can you pass multiple conditions inside `@subset`? e.g. `@subset(:x > 5, :y < 6)`?  
And/or are multiple `@subset` expressions allowed or only one is allowed?

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 26, 2021, 5:16pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/3 "2021-12-26T17:16:22Z")

</div>

Yes multiple conditions are fine, same rules as for the normal `@subset` macro, that’s why I chose this form so it’s hopefully intuitive what’s allowed. Multiple `@subset` macros are not allowed currently, I didn’t think that option would help much.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 26, 2021, 5:21pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/4 "2021-12-26T17:21:01Z")

</div>

> [@jules](#):
>
> I didn’t think that option would help much.

Yes - if multiple conditions are allowed then single `@subset` makes sense. How is `@subset` applied if you pass `GroupedDataFrame` to `transform!` etc.?

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 26, 2021, 5:24pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/5 "2021-12-26T17:24:11Z")

</div>

It’s applied with ungroup = false so that the transform! call also acts on groups afterwards. Then the original grouped dataframe is returned so the result is still grouped.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 26, 2021, 5:36pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/6 "2021-12-26T17:36:52Z")

</div>

> [@jules](#):
>
> Then the original grouped dataframe is returned so the result is still grouped.

Ah - now I see it in the docstring. It is a bit inconsistent as `@transform!` without `@subset` returns the data frame underlying `GroupedDataFrame`.

Why do you prefer to have a different behavior here?

If you wanted to ensure consistency I think the solution would be to add a single check at the beginning of a the macro if an `AbstractDataFrame` or `GroupedDataFrame` is passed and just store there what should be returned at the end.

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 26, 2021, 5:40pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/7 "2021-12-26T17:40:03Z")

</div>

Hm yeah I was unsure about this aspect actually, the difference is that I’m not returning the result of `transform!` because that wouldn’t have all the rows. So if I do ungroup manually, then I should divert the `ungroup = false` option that you could pass to the macro so that it disables my own ungrouping, it wouldn’t do anything in the `transform!` call itself. That also didn’t seem super clear to me, but maybe it’s the better choice?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 26, 2021, 5:46pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/8 "2021-12-26T17:46:28Z")

</div>

I think it could work as follows (pseudocode):

We are in `@transform!` with `@subset` path:

1. if `AbstractDataFrame` store it in some temporary variable `tmp`
2. if `GroupedDataFrme` is passed then:
  - if `ungroup=false` store `parent` of the passed `GroupedDataFrme` in some temporary variable `tmp`
  - if `ungroup=true` store the passed `GroupedDataFrme` in some temporary variable `tmp`

3. perform all the operations you perform (just making sure that proper operations are performed - the returned value does not matter)
4. return `tmp`

This works because in `transform!` and `select!` we know that in the end we should return the original object we were passed.

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 26, 2021, 5:55pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/9 "2021-12-26T17:55:08Z")

</div>

Ok I will consider changing the behavior to this. The tag didn’t go through yet anyway, so there’s no harm.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 26, 2021, 6:52pm UTC](https://discourse.julialang.org/t/ann-dataframemacros-v0-2/73651/10 "2021-12-26T18:52:30Z")

</div>

Sure - pick whatever behavior you think most appropriate.

My reasoning is that using `@subset` should not affect the returned object (except of course the fact that it affects the computation).
