# Remove all entries that occur more than once

**URL:** <https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [February 18, 2022, 6:48am UTC](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672 "2022-02-18T06:48:40Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mo-Gul](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mo-gul/32/33308_2.png) [@Mo-Gul](https://discourse.julialang.org/u/Mo-Gul)\
**Post date:** [February 18, 2022, 6:48am UTC](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672/1 "2022-02-18T06:48:40Z")

</div>

How can I drop all elements in a list that occur more than once? For example when I have

```julia
using DataFrames
firstnames = DataFrame(name = ["Anton", "Jordan", "Jordan"], gender = ["M", "M", "F"])

```

I would like the remaining dataframe to only have the row with the name “Anton”. Unfortunately I can’t use

```julia
unique(firstnames, :name)

```

because it returns the first two rows.

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [February 18, 2022, 7:33am UTC](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672/2 "2022-02-18T07:33:30Z")

</div>

Yeah, `Base.unique` returns distinct values rather than unique ones.

```julia
julia> filter(x->x.nrow ==1, combine(groupby(firstnames, :name), nrow, :gender))
1×3 DataFrame
 Row │ name nrow gender 
     │ String Int64 String 
─────┼───────────────────────
   1 │ Anton 1 M

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [February 18, 2022, 7:34am UTC](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672/3 "2022-02-18T07:34:17Z")

</div>

One way:

```julia
julia> x = filter(:nrow => ==(1), combine(groupby(firstnames, :name), nrow))
1×2 DataFrame
 Row │ name nrow  
     │ String Int64 
─────┼───────────────
   1 │ Anton 1

```

and then subset `firstnames` with this

```julia
julia> firstnames[in(x.name).(firstnames.name), :]
1×2 DataFrame
 Row │ name gender 
     │ String String 
─────┼────────────────
   1 │ Anton M

```

---

<div class="post-metadata">

**Author:** ![Mo-Gul](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mo-gul/32/33308_2.png) [@Mo-Gul](https://discourse.julialang.org/u/Mo-Gul)\
**Post date:** [February 18, 2022, 6:03pm UTC](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672/4 "2022-02-18T18:03:01Z")

</div>

Awesome. Many thanks for the answers. Both are very similar but as a beginner I can follow [jar1’s answer](https://discourse.julialang.org/t/remove-all-entries-that-occur-more-than-once/76672/2) a bit more easily.

Knowing that solution I replace `filter` with `subset` so I can do

```julia
using DataFrames
using Chain

firstnames = DataFrame(name = ["Anton", "Jordan", "Jordan"], gender = ["M", "M", "F"]);
@chain firstnames begin
    groupby(:name)
    combine(:gender, nrow => :nrow)
    subset!(:nrow => ByRow(==(1)))
    select!(Not(:nrow))
end

```

also resulting in

```julia
1×2 DataFrame
 Row │ name gender 
     │ String String 
─────┼────────────────
   1 │ Anton M

```
