# How to select top 5 results per group in dataframe?

**URL:** <https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149>\
**Category:** New to Julia\
**Created:** [September 23, 2020, 6:49pm UTC](https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149 "2020-09-23T18:49:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![conditionality](https://avatars.discourse-cdn.com/v4/letter/c/4af34b/32.png) [@conditionality](https://discourse.julialang.org/u/conditionality)\
**Post date:** [September 23, 2020, 6:49pm UTC](https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149/1 "2020-09-23T18:49:01Z")

</div>

I’m looking for a Julian way of selecting a subset of each group. I have a DataFrame with (among others) two columns, say _name_ and _length_. I want to group over all _names_ and pick the 5 tallest people within each name. I tried this, but it does not return the correct result:

```julia
df = ...
sort!(df, [:length])
df2 = df |> @groupby([:name]) |> @take(5) |> collect
print(DataFrame(df2))

```

Changing collect to DataFrame does not work either. The print will tell me I have a dataframe with as many rows as the initial dataframe **df**. This sort of thing; _taking a df → grouping it → selecting a subset of the rows of the groups → recombining the selected rows into a dataframe_, is something I would assume is a common thing to do.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 23, 2020, 7:19pm UTC](https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149/2 "2020-09-23T19:19:19Z")

</div>

using just DataFrames and `Pipe` you have

```julia
using DataFrames, Pipe
@pipe df |>
    groupby(_, :name) |>
    combine(_) do sdf
        sorted = sort(df, :length)
        first(sorted, 5)
    end

```

---

<div class="post-metadata">

**Author:** ![conditionality](https://avatars.discourse-cdn.com/v4/letter/c/4af34b/32.png) [@conditionality](https://discourse.julialang.org/u/conditionality)\
**Post date:** [September 23, 2020, 7:42pm UTC](https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149/3 "2020-09-23T19:42:10Z")

</div>

Thanks! It feels like the Julian way of doing things is to know which library to use 🙂

Do you mean to use “df” or “sdf” on line 5?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 23, 2020, 7:48pm UTC](https://discourse.julialang.org/t/how-to-select-top-5-results-per-group-in-dataframe/47149/4 "2020-09-23T19:48:42Z")

</div>

sorry, I mean `sdf`. good catch
