# Dataframe Filter

**URL:** <https://discourse.julialang.org/t/dataframe-filter/78491>\
**Category:** New to Julia\
**Tags:** question, dataframes\
**Created:** [March 25, 2022, 10:14pm UTC](https://discourse.julialang.org/t/dataframe-filter/78491 "2022-03-25T22:14:39Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![janilin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/janilin/32/33769_2.png) [@janilin](https://discourse.julialang.org/u/janilin)\
**Post date:** [March 25, 2022, 10:14pm UTC](https://discourse.julialang.org/t/dataframe-filter/78491/1 "2022-03-25T22:14:39Z")

</div>

Hi!  
I was trying to filter data and view the result of some selective columns, but can’t figure out.

I created a dataframe first,

```julia
using DataFrames
df = DataFrame(
    name = ["a", "b", "c", "d"], 
    quiz_1 = [14,15,13,14], 
    quiz_2 = [18,16,16,17], 
    presentation = [7,6,6,8], 
    final = [42,46,47,49])

```

```julia
name	quiz_1	quiz_2	presentation	final
String	Int64	Int64	Int64 Int64
1	a	14 18 7 42
2	b	15 16 6 46
3	c	13 16 6 47
4	d	14 17 8 49

```

In this dataframe, for example, I want to see only `quiz_1` and `quiz_2` columns, who scored more than 45 in the final.

I was able to filter using,

```julia
filter(row -> row.final > 45, df)

```

```julia
	name	quiz_1	quiz_2	presentation	final
String Int64	Int64	Int64 Int64
1	b 15 16 6 46
2	c 13 16 6 47
3	d 14 17 8 49

```

Here, all the columns are shown in the output, but I can’t figure out how to get the result only of `quiz_1` and `quiz_2` columns. Can anyone help me out?

---

<div class="post-metadata">

**Author:** ![AndiMD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andimd/32/6531_2.png) [@AndiMD](https://discourse.julialang.org/u/AndiMD)\
**Post date:** [March 25, 2022, 10:37pm UTC](https://discourse.julialang.org/t/dataframe-filter/78491/2 "2022-03-25T22:37:53Z")

</div>

Hi, these are some basic options:  
`df[df.final.>45,[:quiz_1,:quiz_2]]`  
`filter(row -> row.final > 45, df)[:,[:quiz_1,:quiz_2]]`  
`DataFrame(i[[:quiz_1,:quiz_2]] for i in eachrow(df) if i.final>45)`

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 26, 2022, 12:02am UTC](https://discourse.julialang.org/t/dataframe-filter/78491/3 "2022-03-26T00:02:35Z")

</div>

You can also use DataFramesMeta.jl for dplyr-like syntax

```julia
@rsubset df :final > 45 # r stands for row-wise

```

---

<div class="post-metadata">

**Author:** ![janilin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/janilin/32/33769_2.png) [@janilin](https://discourse.julialang.org/u/janilin)\
**Post date:** [March 26, 2022, 5:27am UTC](https://discourse.julialang.org/t/dataframe-filter/78491/4 "2022-03-26T05:27:57Z")

</div>

Thanks a lot for sharing the solution!

---

<div class="post-metadata">

**Author:** ![janilin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/janilin/32/33769_2.png) [@janilin](https://discourse.julialang.org/u/janilin)\
**Post date:** [March 26, 2022, 5:32am UTC](https://discourse.julialang.org/t/dataframe-filter/78491/5 "2022-03-26T05:32:05Z")

</div>

Thanks for the suggestion! Going through the documentation, pretty dplyr-like syntax indeed.

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [March 26, 2022, 7:17am UTC](https://discourse.julialang.org/t/dataframe-filter/78491/6 "2022-03-26T07:17:41Z")

</div>

While the effect is minimal in this example, you can gain a bit of efficiency by passing only the column(s) you’re interested in to filter i.e.

```julia
julia> @btime filter(row -> row.final>45, df)
  18.120 μs (25 allocations: 1.88 KiB)

```

vs

```julia
julia> @btime filter(:final => f-> f>45, df)
  15.324 μs (24 allocations: 1.81 KiB)

```

for multiple

```julia
julia> @btime filter(row -> row.name=="b" && row.final>45, df)
  19.188 μs (25 allocations: 1.80 KiB)

```

vs

```julia
julia> @btime filter([:name, :final] => (n,f)-> n=="b" && f>45, df)
  18.529 μs (29 allocations: 1.98 KiB)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [March 26, 2022, 7:38am UTC](https://discourse.julialang.org/t/dataframe-filter/78491/7 "2022-03-26T07:38:39Z")

</div>

FWIW, for the specific OP’s dataframe example the type of solution proposed by AndiMD seems to run faster than all other alternatives shown:

```julia
@btime $df[$df.final .> 45, [:quiz_1,:quiz_2]] # 1.320 μs (26 allocs: 1.7 KiB)
@btime $df[$df.name .== "b" .&& $df.final .> 45,[:quiz_1,:quiz_2]] # 1.610 μs (30 allocs: 1.8 KiB)

```
