# Find unique row in DataFrame

**URL:** <https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945>\
**Category:** General Usage\
**Created:** [May 17, 2018, 12:53am UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945 "2018-05-17T00:53:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [May 17, 2018, 12:53am UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/1 "2018-05-17T00:53:59Z")

</div>

Hi All,

I am finding unique rows in a DataFrame and get number of repeat of unique rows, I could get by this:

```nohighlight
t=DataFrame(a=rand(1:5,20),b=rand([:x,:y,:z],20))
ut=unique(t)

Int[countnz([t[j,:]==ut[i,:] for j in 1:size(t,1)]) for i in 1:size(ut,1)]

```

but I am wondering if there are more clean way, such as

```nohighlight
Int[countnz(t.==r) for r in t]
or
Int[countnz(t.==r) for r in eachrow(t)]

```

right now, none of the above works, however it came to me as a natural way to broadcast each row and compare.

any thoughts?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [May 17, 2018, 1:43am UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/2 "2018-05-17T01:43:37Z")

</div>

You can do this with a grouped operation

```julia
# you can group with a vector of symbols, so use all names in the DataFrame
by(t, names(t)) do # do is an easy way to do an anonymous function
       DataFrame(m = length(d[:a])) # just choose any column
end

```

With `DataFramesMeta` and `Lazy`, which is closest to R’s chaining imo, you can do

```julia
using DataFramesMeta, Lazy
t = @> t begin
    @by(names(t), n_copies = length(:a))
end

```

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [May 17, 2018, 4:25am UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/3 "2018-05-17T04:25:32Z")

</div>

@pdeffebach thanks for the tip, split-apply-combine is more elegant.

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [May 17, 2018, 5:44am UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/4 "2018-05-17T05:44:16Z")

</div>

@pdeffebach  
btw, how do I get the indices of each row in grouped DataFrame, so that I can index back to the original ungrouped DataFrame. like this:

```nohighlight
[find([t[j,:]==ut[i,:] for j in 1:size(t,1)]) for i in 1:size(ut,1)]

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [May 17, 2018, 1:50pm UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/5 "2018-05-17T13:50:30Z")

</div>

Comparing rows across DataFrames is less straightforward than I thought, but it might be easier on `master`, I’m not sure.

Here’s the `DataFramesMeta` way:

```julia
t[:rownum] = [i for i in nrow(t)]
ut = @> t begin
     @by([:a,:b], places = [[i for i in :rownum]])
end

```

---

<div class="post-metadata">

**Author:** ![babaq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/babaq/32/891_2.png) [@babaq](https://discourse.julialang.org/u/babaq)\
**Post date:** [May 17, 2018, 9:29pm UTC](https://discourse.julialang.org/t/find-unique-row-in-dataframe/10945/6 "2018-05-17T21:29:20Z")

</div>

@pdeffebach thanks, I did it this way:

```nohighlight
cols=names(t)
t[:I] = 1:nrow(t)
by(t, cols,g->DataFrame(n=size(g,1), i=[g[:i]]))

```
