# Deterministic Clustering in Julia?

**URL:** https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068
**Category:** Machine Learning
**Tags:** clustering
**Created:** [June 25, 2020, 9:05pm UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068 "2020-06-25T21:05:41Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)
#### Post date: [June 25, 2020, 9:05pm UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068/1 "2020-06-25T21:05:41Z")

</div>

Are there any deterministic clustering methods implemented in Julia? For example:

```julia
using Clustering
using DataFrames

df = DataFrame(subject=1:10, height=rand(10), weight=rand(10))

julia> groups = kmeans(Matrix(df[:, 2:3])', 2).assignments
10-element Array{Int64,1}:
 1
 1
 2
 1
 1
 1
 2
 1
 1
 1

julia> groups = kmeans(Matrix(df[:, 2:3])', 2).assignments
10-element Array{Int64,1}:
 1
 1
 1
 2
 2
 2
 1
 1
 1
 1

```

Is there a clustering algorithm in Julia that will result in the same clusters for this problem each time? Or is there a way to set the seed via `Random.seed!()` that would result in the same clusters with k-means?

Any recommendations as to how to approach a problem like this where a deterministic outcome is very important would be much appreciated!

---

<div class="post-metadata">

### Author: ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)
#### Post date: [June 26, 2020, 2:07am UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068/2 "2020-06-26T02:07:58Z")

</div>

So I looked at the Clustering.jl docs… and found that you can control the Seeding algorithm:  
[https://juliastats.org/Clustering.jl/stable/kmeans.html](https://juliastats.org/Clustering.jl/stable/kmeans.html)  
[https://juliastats.org/Clustering.jl/stable/init.html#Seeding-1](https://juliastats.org/Clustering.jl/stable/init.html#Seeding-1)

You can provide your own vector of k indices of points to use for the seeds. So if you have some way to assign cluster seeds that way, it should be deterministic after that.

Have you tried setting the random seed methodology?

---

<div class="post-metadata">

### Author: ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)
#### Post date: [June 26, 2020, 3:26am UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068/3 "2020-06-26T03:26:57Z")

</div>

I think that hdbscan is deterministic  
I have a wrapper around the Python package here  
[https://github.com/baggepinnen/HDBSCAN.jl](https://github.com/baggepinnen/HDBSCAN.jl)

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [June 26, 2020, 6:05am UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068/4 "2020-06-26T06:05:32Z")

</div>

Sure there is 🙂

There is also kmeans in [BetaML](https://sylvaticus.github.io/BetaML.jl/dev/Clustering.html#BetaML.Clustering.kmeans-Tuple%7BAny,Any%7D) (disclaimer: I am the author), chose as `initStrategy` parameter `grid` (default) or `given`.

In `given` you provide your own init points, in `grid` it scans the input space and starts at regular grid intervals:

```julia
julia> (dataAllocations,clusterMeans) = kmeans([1 10.5;1.5 10.8; 1.8 8; 1.7 15; 3.2 40; 3.6 32; 3.3 38; 5.1 -2.3; 5.2 -2.4],3)

```

Also em by default is deterministic, although initiating the mixtures with the result of kmeans is much a better init approach.

(If kmeans doesn’t converge, try master… I just committed yesterday a correction for a corner case)

---

<div class="post-metadata">

### Author: ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)
#### Post date: [June 26, 2020, 6:54am UTC](https://discourse.julialang.org/t/deterministic-clustering-in-julia/42068/5 "2020-06-26T06:54:41Z")

</div>

You can also use [ParallelKMeans.jl](https://github.com/PyDataBlog/ParallelKMeans.jl). All algorithms support `rng` parameter, so you can call it like that:

```julia
using Random, ParallelKMeans

rng = Random.seed!(2020)
kmeans(X, 10; rng = rng)

```

or you can use [StableRNGs.jl](https://github.com/rfourquet/StableRNGs.jl) if you want to have the same results across all julia versions

```julia
using StableRNGs, ParallelKMeans

rng = StableRNG(2020)
kmeans(X, 10; rng = rng)

```

Also, for `Lloyd()`, `Hammerly()` and `Elkan()` algorithms, random generation is used only for initial seeding, so if you can prepare `init`, you can get reproducible results with

```julia
kmeans(Hammerly(), X, 10; init = init)

```

In addition, this implementation is really fast, which can be useful.
