# R function \`duplicated\`

**URL:** <https://discourse.julialang.org/t/r-function-duplicated/91707>\
**Category:** General Usage\
**Tags:** function\
**Created:** [December 15, 2022, 5:29pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707 "2022-12-15T17:29:40Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 5:29pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/1 "2022-12-15T17:29:40Z")

</div>

Hello,

For a vector `v`, the R function `duplicated` works as follows. The vector `duplicated(v)` has the same length as `v`, and its `i`-th element is `false` if and only if `v[i]` is the first occurence of `v[i]` in `v`. For example `duplicated([1, 2, 1, 3, 2]) = [false, false, true, false, true]`. I implemented it as follows in Julia:

```julia
function duplicated(x)
    out = fill(false, length(x))
    for i in 1:(length(x)-1)
        if !out[i]
            out[i .+ findall(x[i] .== x[(i+1):length(x)])] .= true
        end
    end
    return out
end

```

Can we improve it?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [December 15, 2022, 5:35pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/2 "2022-12-15T17:35:56Z")

</div>

this feels like an intermediate masking kind of thing that’s more useful for R/Python than Julia, what are you gonna do with this vector?

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [December 15, 2022, 5:35pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/3 "2022-12-15T17:35:57Z")

</div>

Something like this (untested code) will make it run in O(n log n) instead of O(n^2):

> [@\_stla](#):
>
> ```julia
> function duplicated(x)
> out = fill(false, length(x))
> seen = Set{eltype(x)}()
> for (i,value) in enumerate(x)
> out[i] = value in seen
> push!(seen,value)
> end
> return out
> end
> 
> ```

---

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 5:44pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/4 "2022-12-15T17:44:27Z")

</div>

To remove the duplicates of a vector you can do `v[!duplicated(v)]`. Well, in this case this is equivalent to `unique(v)`, but this can be used for another vector: `x[!duplicated(v)]` (useful for example `v = score.(x)` for a function `score`).

---

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 5:46pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/5 "2022-12-15T17:46:22Z")

</div>

Thanks. Are you sure it is better? With my function, elements marked as duplicates are not tested a second time in the next tests.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [December 15, 2022, 5:53pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/6 "2022-12-15T17:53:01Z")

</div>

> [@\_stla](#):
>
> `x[!duplicated(v)]` (useful for example `v = score.(x)` for a function `score`).

```julia
unique(score, x)

```

---

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 5:55pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/7 "2022-12-15T17:55:53Z")

</div>

Didn’t know that, thanks. But I use it for removing the rows of a matrix which have the same “score”: `x[!duplicated([score(row) for row in eachrow(x)]), :]`.

---

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 6:02pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/8 "2022-12-15T18:02:15Z")

</div>

Ah I see, I could use `unique(score, collect(eachrow(x)))` instead and reconstruct a matrix with these rows. But this should be less efficient no?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 15, 2022, 6:23pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/9 "2022-12-15T18:23:12Z")

</div>

> [@\_stla](#):
>
> To remove the duplicates of a vector you can do `v[!duplicated(v)]`.

Couldn’t you do

```julia
i = unique(j -> v[j], eachindex(v))
v[i]

```

for this purpose? This gives you an index array that you could re-use to extract corresponding slices of other arrays too.

See also the discussions at [Return index vectors from unique · Issue #1845 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/1845) and  
[Is there a function similar to numpy unique with inverse? - #6 by stevengj](https://discourse.julialang.org/t/is-there-a-function-similar-to-numpy-unique-with-inverse/80949/6)

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [December 15, 2022, 6:50pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/10 "2022-12-15T18:50:20Z")

</div>

> [@\_stla](#):
>
> I could use `unique(score, collect(eachrow(x)))`

```julia
unique!.(score, eachrow(x))

```

this modifies `x` in-place instead of making copies

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 15, 2022, 7:37pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/11 "2022-12-15T19:37:04Z")

</div>

In DataFrames.jl you can do:

```julia
julia> df = DataFrame(x = rand(1:10^6, 10^6));

julia> @time nonunique(df, :x);
  0.034752 seconds (52 allocations: 24.585 MiB)

```

(which gives you a `Bool` vector exactly as you want and is slightly faster than the `unique` version that @stevengj proposed)

---

<div class="post-metadata">

**Author:** ![\_stla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/_stla/32/35613_2.png) [@\_stla](https://discourse.julialang.org/u/_stla)\
**Post date:** [December 15, 2022, 7:37pm UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/12 "2022-12-15T19:37:52Z")

</div>

Yes, I finally had the same idea. More generally `unique(j -> score(v[j]), eachindex(v))`.

---

<div class="post-metadata">

**Author:** ![skleinbo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skleinbo/32/36080_2.png) [@skleinbo](https://discourse.julialang.org/u/skleinbo)\
**Post date:** [December 16, 2022, 5:57am UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/13 "2022-12-16T05:57:19Z")

</div>

> [@Jeff\_Emanuel](#):
>
> will make it run in O(n log n) instead of O(n^2):

Wouldn’t it be rather \mathcal{O}(n\cdot k) with n the length of the input and k the number of its unique elements? Which is n^2 if all elements are unique.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [December 16, 2022, 6:34am UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/14 "2022-12-16T06:34:40Z")

</div>

You have `n` to loop over the values. For each value you do a set lookup and a possible set insertion (and setting the boolean value in `out`, which is clearly constant time). I mentioned `log n` thinking that a set might be a binary tree. If instead a set is implemented as a hash (likely), then insertion and lookup could be constant time, so its possibly (probable) just O(n). I’m not at a computer so it’s not convenient to lookup the set implementation details.

---

<div class="post-metadata">

**Author:** ![skleinbo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skleinbo/32/36080_2.png) [@skleinbo](https://discourse.julialang.org/u/skleinbo)\
**Post date:** [December 16, 2022, 7:02am UTC](https://discourse.julialang.org/t/r-function-duplicated/91707/15 "2022-12-16T07:02:16Z")

</div>

Ah, you’re of course totally right. Sorry, that wasn’t very clever of me🤦‍♂️. It is a hasmap by the way, so `O(n)` it is!
