# A nice use case for DataFrames.jl - flexible dedup

**URL:** <https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725>\
**Category:** General Usage\
**Tags:** dataframes, tables, splitapplycombine\
**Created:** [July 16, 2021, 1:59am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725 "2021-07-16T01:59:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [July 16, 2021, 1:59am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725/1 "2021-07-16T01:59:08Z")

</div>

I had this problem where I need to dedup a dataframe using a column, but there is a timestamp column for all the entries. If two entries have the same data, but different timestamp, I would want ot keep only the one with the latest timestamp. After this dedup there shouldn’t any dups in the dedup column.

Obviously `unique` wouldn’t work due to timestamp and other columns. This is how i solved it with DataFrames.jl

```julia
using DataFrames

data=DataFrame(dedup = rand(1:8, 1000), timestamp = rand(1:8, 1000), othervals = rand(1000))

using Chain
@chain data begin
  groupby(:dedup)
  combine(subdf -> begin
    sort(subdf, :timestamp, rev=true)[1, :] # could've used partialsort but this is easier to read
  end)
end

```

which I thought was a very nice and readable way to do it.

How do you guys do it? Any examples from other languages? I think data.table nor dplyr is as nice. And don’t get me started on pandas or spark… Can’t rule out i am just bad at pandas and spark though 🙂

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 16, 2021, 7:04am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725/2 "2021-07-16T07:04:07Z")

</div>

> [@xiaodai](#):
>
> How do you guys do it?

Couldn’t be easier with `Base` Julia `Table`s as well (:

```julia
using SplitApplyCombine, Tables, DataPipes

tbl = (dedup = rand(1:8, 1000), timestamp = rand(1:8, 1000), othervals = rand(1000)) |> rowtable

# closest to your solution:
@p begin
	tbl
	group(_.dedup)
	map() do subtbl
		@p subtbl |> sort(by=_.timestamp) |> last
	end
	rowtable
end

# write the short lambda inline:
@p begin
	tbl
	group(_.dedup)
	map(sort(_, by=x->x.timestamp) |> last)
	rowtable
end

# same without macros, less clear:
map(
	subtbl -> sort(subtbl, by=x->x.timestamp)[end],
	group(x -> x.dedup, tbl)
) |> rowtable

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [July 16, 2021, 7:21am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725/3 "2021-07-16T07:21:30Z")

</div>

Still essentially a group by approach. Wondering if there are non group by approaches

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 16, 2021, 7:26am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725/4 "2021-07-16T07:26:12Z")

</div>

Of course, one could just use `unique`:

```julia
@p begin
	tbl
	sort(by=_.timestamp, rev=true)
	unique(_.dedup)
end

```

Its docs say:

> Return an array containing only the unique elements of collection `itr` , as determined by `isequal`, **in the order that the first of each set of equivalent elements originally appears**.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 16, 2021, 7:53am UTC](https://discourse.julialang.org/t/a-nice-use-case-for-dataframes-jl-flexible-dedup/64725/5 "2021-07-16T07:53:34Z")

</div>

If you need performance `argmin` instead of `sort` would be faster.
