# ByRow subset vs filter performance

**URL:** <https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022>\
**Category:** Data\
**Tags:** dataframes\
**Created:** [August 19, 2022, 4:39pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022 "2022-08-19T16:39:01Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 19, 2022, 4:39pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/1 "2022-08-19T16:39:01Z")

</div>

I just compared `subset` (using a `:col => ByRow(==(0))` operation) vs `filter` (using a `:col => ==(0)` operation) on a DataFrame. In both cases, I set `view=true` and I see significantly less allocations when using `filter` vs when using `subset`. Would you agree that in these cases, `filter` is more performant than the `subset` operation?

Thanks!

Edit, here is a MWE:

```julia
df=DataFrame(A = rand(10))
f1(df) = @time subset(df, :A => ByRow(<=(0.5)), view=true)
f2(df) = @time filter(:A => <=(0.5), df, view=true)

f1(df);
  0.000091 seconds (162 allocations: 8.391 KiB)

f2(df);
  0.000011 seconds (8 allocations: 320 bytes)

```

---

<div class="post-metadata">

**Author:** ![digital\_carver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/digital_carver/32/33818_2.png) [@digital\_carver](https://discourse.julialang.org/u/digital_carver)\
**Post date:** [August 19, 2022, 5:29pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/2 "2022-08-19T17:29:25Z")

</div>

A more typical benchmark using `@btime` seems to support this result too (with logical indexing leading to slightly better performance than `filter` despite having more allocations):

```julia
julia> begin
         df = DataFrame(A = rand(10000))
         f1(df) = subset(df, :A => ByRow(<=(0.5)), view = true)
         f2(df) = filter(:A => <=(0.5), df, view = true)
         f3(df) = subset(df, :A => a -> a .<= 0.5, view = true)
         f4(df) = @view df[df.A .<= 0.5, :] 
         @btime f1(df)
         @btime f2(df)
         @btime f3(df)
         @btime f4(df)
       end;
  602.734 ms (464204 allocations: 8.92 MiB)
  379.465 ms (298631 allocations: 4.98 MiB)
  924.066 ms (640950 allocations: 11.75 MiB)
  372.004 ms (328600 allocations: 6.05 MiB)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 19, 2022, 6:46pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/3 "2022-08-19T18:46:25Z")

</div>

> [@hdavid16](#):
>
> Would you agree that in these cases, `filter` is more performant than the `subset` operation?

This is expected in this case.

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [August 19, 2022, 7:55pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/4 "2022-08-19T19:55:29Z")

</div>

Could you explain why should filter be faster than subset?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 19, 2022, 8:06pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/5 "2022-08-19T20:06:44Z")

</div>

In short: because `filter` accepts only one condition + it works rowwise, so it has much simpler logic internally. `subset` allows passing multiple conditions + it works on whole columns. (if you want more details it is best to check the source code to see the differences in implementation)

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 19, 2022, 8:13pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/6 "2022-08-19T20:13:57Z")

</div>

Thanks for pointing this out @bkamins. In general, could we say that if we are doing row-wise filtering, even on multiple conditions (joined with `&&` operators), we should expect `filter` to outperform `subset`?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 19, 2022, 8:39pm UTC](https://discourse.julialang.org/t/byrow-subset-vs-filter-performance/86022/7 "2022-08-19T20:39:15Z")

</div>

yes. but you need to hardcode `&&` in the predicate. Of course this assumes you use the  
`filter(cols => predicate, df)` style.
