# Rank with ties

**URL:** <https://discourse.julialang.org/t/rank-with-ties/84512>\
**Category:** New to Julia\
**Created:** [July 20, 2022, 6:52am UTC](https://discourse.julialang.org/t/rank-with-ties/84512 "2022-07-20T06:52:47Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [July 20, 2022, 6:52am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/1 "2022-07-20T06:52:47Z")

</div>

Suppose I have a vector containing the following integers: [1 1 1 2 3 5 5 9]

Now I wish to create a vector that ranks the integers in the following way:

rank = [1 1 1 2 3 4 4 5]

The rank essentially answers the question for each integer “number of unique integers larger + 1” but implementing that is causing some problems. Maybe there is a function already?

This is my attempt, but how do I broadcast it to the vector at once?

length(unique(A[A .\< A[8]])) + 1

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 20, 2022, 7:03am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/2 "2022-07-20T07:03:41Z")

</div>

See

> **[GitHub - oheil/NormalizeQuantiles.jl: NormalizeQuantiles.jl implements...](https://github.com/oheil/NormalizeQuantiles.jl)**
>
> NormalizeQuantiles.jl implements quantile normalization - GitHub - oheil/NormalizeQuantiles.jl: NormalizeQuantiles.jl implements quantile normalization

function `sampleRanks`: [GitHub - oheil/NormalizeQuantiles.jl: NormalizeQuantiles.jl implements quantile normalization](https://github.com/oheil/NormalizeQuantiles.jl#usage-examples-sampleranks)

---

<div class="post-metadata">

**Author:** ![SteffenPL](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steffenpl/32/206270_2.png) [@SteffenPL](https://discourse.julialang.org/u/SteffenPL)\
**Post date:** [July 20, 2022, 7:14am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/3 "2022-07-20T07:14:09Z")

</div>

An elementary way would be

```julia
x = [1,1,1,2,3,5,5,9]
xu = unique(x)
ranks = [count( xu .< z ) + 1 for z in x]

```

(Of course, @oheil’s package does more than that, i.e. it also takes care of edge cases and supports matrix inputs etc.)

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 20, 2022, 7:20am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/4 "2022-07-20T07:20:09Z")

</div>

> [@SteffenPL](#):
>
> (Of course, @oheil’s package does more than that, i.e. it also takes care of edge cases and supports matrix inputs etc.)

Thats why it’s typically slower…

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 20, 2022, 8:22am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/5 "2022-07-20T08:22:52Z")

</div>

`StatsBase.denserank` seems to be exactly what you need.

---

<div class="post-metadata">

**Author:** ![maxkapur](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxkapur/32/21208_2.png) [@maxkapur](https://discourse.julialang.org/u/maxkapur)\
**Post date:** [July 21, 2022, 12:43am UTC](https://discourse.julialang.org/t/rank-with-ties/84512/6 "2022-07-21T00:43:43Z")

</div>

@SteffenPL’s function (and any implementation based on `count`) will have worst-case O(n^2) complexity (when all the entries of `x` are different).

But by sorting `x` out the outset, it is possible to do this in O(n \log n)-time using something like this:

```julia
function denserank(x::Vector)
    p = sortperm(x) # to recover sort later
    permute!(x, p)
    res = ones(Int, length(x))
    currentrank = 1
    for i in 2:length(x)
        if x[i-1] < x[i]
            currentrank += 1
        end
        res[i] = currentrank
    end
    invpermute!(res, p)
end

```

```julia
julia> x = [1, 6, 3, 3, 7, 1, 1]; denserank(x)
7-element Vector{Int64}:
 1
 3
 2
 2
 4
 1
 1

```

This is the approach used by `StatsBase.denserank`, albeit as the source code [here](https://github.com/JuliaStats/StatsBase.jl/blob/77590cee52eb16c960710789c9ef323dd551d580/src/ranking.jl#L93) reveals, there are a few layers of abstraction in between so it can support more than just `Vector` input.

In summary, you should either use `StatsBase.denserank`, or if you roll your own function, base it on a sort algorithm instead of than `count`.
