# DataFramesMeta conditionals

**URL:** <https://discourse.julialang.org/t/dataframesmeta-conditionals/10057>\
**Category:** General Usage\
**Created:** [March 29, 2018, 7:59pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057 "2018-03-29T19:59:59Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [March 29, 2018, 7:59pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/1 "2018-03-29T19:59:59Z")

</div>

I often forget to put the dot when writing conditions. I wonder why we couldn’t just do it without and automate it inside DataFramesMeta with a `@.` macro. Any thoughts?

So, rather than writing this:

```julia
@where(df, :x .> 1)

```

I can do just this:

```julia
@where(df, :x > 1)

```

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [March 29, 2018, 8:02pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/2 "2018-03-29T20:02:11Z")

</div>

I rather like the idea that the symbols are a simple drop in for the actual columns. It might save you a little bit of typing to auto-broadcast or whatever, but at the cost of making the overall semantics more confusing. Considering that people are free to put whatever function they want into the `@where` statements, including functions they define themselves, I’d think that complicating the semantics would be a bad idea.

---

<div class="post-metadata">

**Author:** ![tshort](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tshort/32/43_2.png) [@tshort](https://discourse.julialang.org/u/tshort)\
**Post date:** [March 29, 2018, 8:34pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/3 "2018-03-29T20:34:39Z")

</div>

See also the following issue: [https://github.com/JuliaStats/DataFramesMeta.jl/issues/39](https://github.com/JuliaStats/DataFramesMeta.jl/issues/39)

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [March 29, 2018, 8:44pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/4 "2018-03-29T20:44:12Z")

</div>

JuliaDBMeta distinguishes [row-wise macros](https://piever.github.io/JuliaDBMeta.jl/latest/row_macros.html) (`@map`, `@where`, `@transform`) where your are iterating through rows and symbols correspond to a given field and [column-wise macros](https://piever.github.io/JuliaDBMeta.jl/latest/column_macros.html) (`@with`, `@where_vec`, `@transform_vec`) where symbols correspond to columns and you will often need to use dot broadcasting in combination with the latter.

Note that both versions are required, as for example one may want:

`@where_vec(df, :a .> mean(:a))`

which can’t be achieved row by row.

I wonder whether DataFramesMeta could implement a similar strategy. I’m not sure how easy it is to implement row-wise macros efficiently due to type stability issues with DataFrames, but maybe there are ways around that.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [March 29, 2018, 8:49pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/5 "2018-03-29T20:49:15Z")

</div>

In R dplyr, `filter(a > mean(a))` works. It is much less verbose.

---

<div class="post-metadata">

**Author:** ![pasha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pasha/32/3319_2.png) [@pasha](https://discourse.julialang.org/u/pasha)\
**Post date:** [March 29, 2018, 8:52pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/6 "2018-03-29T20:52:13Z")

</div>

It seems to me that unless we have a `@byrow!` in play, we’re always in column space. I don’t see where the promotion of `>` to `.>` would cause confusion; I don’t know what someone would _mean_ in this context with a non-broadcast `>`. It may be technically hard to achieve, but I support making it all `@.` if possible.

Addendum: This reminds me of the difference between `mean` and `pmean`, likewise `max` and `pmax`, `min` and `pmin`. That is a tricky place in R/dplyr, [it’s basically a local max/min/mean calculation](https://dennisphdblog.wordpress.com/2009/07/24/r-command-of-the-week-pmax-and-pmin/). So this nuance could be confused by the above syntax… maybe.

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [March 29, 2018, 9:01pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/7 "2018-03-29T21:01:36Z")

</div>

The only differences I see between `filter(a > mean(a))` and `@where_vec(:a .> mean(:a))` (JuliaDBMeta, just like dpyr, has a curried version) is the use of symbols to refer to columns and of dot broadcasting for element-wise comparison. I’m really not sure how one can avoid using symbols and just put variable names. Dot broadcasting is necessary because, if we are taking `mean(a)` it means `a` is a vector and thus we need to compare element-wise.

The example is in my view interesting because it shows a case where automatic dot broadcasting would not work with `@where` in DataFrames as one would get `v .> mean.(v)` which is not the correct thing.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [March 29, 2018, 9:49pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/8 "2018-03-29T21:49:38Z")

</div>

a is a array, mean(a) is a scalar, when compare a and mean(a), I guess that the scalar is automatically converted to a array with the same length. I do not think there should be any confusion. I have used dplyr for my daily work for several years, and it works all well.

Pandas and data.table are a little bit more verbose than dplyr. It is a pain to do data wrangling in Matlab, and Julia seems to use Matlab style syntax for data manipulation. I think such style is great for writing numerical stuff, but for data manipulation, it might be easier to follow R and Python style.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 29, 2018, 11:05pm UTC](https://discourse.julialang.org/t/dataframesmeta-conditionals/10057/9 "2018-03-29T23:05:32Z")

</div>

What about this use case?

```julia
# keep only observations above the mean of income

@where(df, :income .> mean(:income)

```

As opposed to

```julia
m = mean(df[:income])
@where(df, :income > m)

```
