# How to filter single subject from dataframe

**URL:** <https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [April 22, 2021, 11:02am UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793 "2021-04-22T11:02:51Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 22, 2021, 11:02am UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/1 "2021-04-22T11:02:51Z")

</div>

I am started learning julia recently. I am working with a dataset where i want to filter all the rows of single ID in a data frame.  
can somebody help me to learn this?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [April 22, 2021, 11:27am UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/2 "2021-04-22T11:27:58Z")

</div>

There is no way to do that in an “easy” way right now in DataFrames.

If your `ID` are unique, you can do

```julia
df[findfirst(==("idnum"), df.ID), :] |> first

```

- `findfirst(==("idnum"), df.ID)` finds the furst occurance of `"idnum"` in `df.ID`
- The indexing will return a `DataFrame` with one row. You call `first` to get a `DataFrameRow` which is a nicer object for this purpose.

---

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 22, 2021, 4:18pm UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/3 "2021-04-22T16:18:12Z")

</div>

Thank you

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [April 23, 2021, 8:53am UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/4 "2021-04-23T08:53:21Z")

</div>

Trying to understand the question and how different it is from the following MWE:

```julia
using DataFrames

team = DataFrame(ID=[1,2,3,4,2,4],name=["John","Jane","Jim","Joe","Jay","Julia"])

team[team.ID .== 4, :]

```

which for the input:

```julia
│ Row │ ID │ name │
│ │ Int64 │ String │
├─────┼───────┼────────┤
│ 1 │ 1 │ John │
│ 2 │ 2 │ Jane │
│ 3 │ 3 │ Jim │
│ 4 │ 4 │ Joe │
│ 5 │ 2 │ Jay │
│ 6 │ 4 │ Julia │

# produces the filtered result by team ID number 4:
│ Row │ ID │ name │
│ │ Int64 │ String │
├─────┼───────┼────────┤
│ 1 │ 4 │ Joe │
│ 2 │ 4 │ Julia │ 

```

Or should the output sought consist of only the teams with one member each (i.e., 1 and 3):

```julia
using StatsBase

dic = countmap(team.ID)
ix = [dic[x]==1 for x in team.ID]
team[ix,:] 

│ Row │ ID │ name │
│ │ Int64 │ String │
├─────┼───────┼────────┤
│ 1 │ 1 │ John │
│ 2 │ 3 │ Jim │

```

Thanks.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [April 23, 2021, 8:03pm UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/5 "2021-04-23T20:03:08Z")

</div>

I think it’s assumed that IDs uniquely identify observations here.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 23, 2021, 8:17pm UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/6 "2021-04-23T20:17:19Z")

</div>

> [@pdeffebach](#):
>
> `df[findfirst(==("idnum"), df.ID), :]`

this will already return `DataFrameRow`

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 23, 2021, 8:18pm UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/7 "2021-04-23T20:18:50Z")

</div>

> [@rafael.guerra](#):
>
> `team[team.ID .== 4, :]`

and to make sure you have only one row just write:

```julia
only(team[team.ID .== 4, :])

```

or

```julia
only(filter(:ID => ==(4), team))

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [April 23, 2021, 9:20pm UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/8 "2021-04-23T21:20:16Z")

</div>

@bkamins, thanks for the feedback. That command throws an error when run for the non-unique ID-rows (2 and 4):

`ERROR: ArgumentError: data frame must contain exactly 1 row`

Why not throwing an empty data frame instead (_if such creature exists_)?

_ **NB:** the MWE above using `countmap()` outputs only the ID rows with multiplicity=1. May it be written more simply within DataFrames framework?_

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 24, 2021, 12:00am UTC](https://discourse.julialang.org/t/how-to-filter-single-subject-from-dataframe/59793/9 "2021-04-24T00:00:08Z")

</div>

I used `only` to show that you can use it to check you get a result that has 1 row (as this was suggested by @pdeffebach above).

Empty data frame `DataFrame()` exists, but this is not how `only` is defined in Julia Base (you can check its doctsting to find the contract it guarantees).

> _ **NB:** the MWE above using `countmap()` outputs only the ID rows with multiplicity=1. May it be written more simply within DataFrames framework?_

```julia
julia> DataFrame(filter(sdf -> nrow(sdf)==1, groupby(team, :ID)))
2×2 DataFrame
 Row │ ID name   
     │ Int64 String 
─────┼───────────────
   1 │ 1 John
   2 │ 3 Jim

```

or

```julia
julia> combine(groupby(team, :ID), sdf -> nrow(sdf) == 1 ? sdf : DataFrame())
2×2 DataFrame
 Row │ ID name   
     │ Int64 String 
─────┼───────────────
   1 │ 1 John
   2 │ 3 Jim

```
