# Looking for Julia implementations of online clustering algorithms

**URL:** https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649
**Category:** Statistics
**Tags:** package, clustering
**Created:** [April 30, 2024, 1:20pm UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649 "2024-04-30T13:20:49Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![vvbond](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vvbond/32/10105_2.png) [@vvbond](https://discourse.julialang.org/u/vvbond)
#### Post date: [April 30, 2024, 1:20pm UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649/1 "2024-04-30T13:20:49Z")

</div>

… to no avail.

The use case is quite common: real time clustering of streamed data.  
I envision a fusion between functionalities of the excellent OnlineStats.jl and Clustering.jl, but so far couldn’t find any ready implementations.

I’d appreciate any pointer to relevant Julia project(s). Thanks!

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [April 30, 2024, 6:55pm UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649/2 "2024-04-30T18:55:53Z")

</div>

Probably not efficient for very large datasets (that I guess is the need for online training), but my [BetaML.GaussianMixtureClustering](https://sylvaticus.github.io/BetaML.jl/dev/GMM.html#BetaML.GMM.GaussianMixtureClusterer) model has _some_ support for online fitting:

“Online fitting (re-fitting with new data) is supported by setting the old learned mixtrures as the starting values”

---

<div class="post-metadata">

### Author: ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)
#### Post date: [April 30, 2024, 6:59pm UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649/3 "2024-04-30T18:59:49Z")

</div>

Give `OnlineStats.Kmeans` a try:

> **[API · OnlineStats Documentation](https://joshday.github.io/OnlineStats.jl/latest/api/#OnlineStats.KMeans)**
>
> Documentation for OnlineStats Documentation.

---

<div class="post-metadata">

### Author: ![vvbond](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vvbond/32/10105_2.png) [@vvbond](https://discourse.julialang.org/u/vvbond)
#### Post date: [May 1, 2024, 10:26am UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649/4 "2024-05-01T10:26:12Z")

</div>

Yeah, the usual problem with k-Means is the need to specify the number of clusters beforehand. In my problem, the data stream is expected to bifurcate into several clusters which then can go stale, i.e., disappear for a while, reappear, etc.

OnlineStats.jl seems to offer many building blocks useful to tackle the problem and I will give it a try, but I hoped to find something ready made. Perhaps, something like Python’s [DenStream - River](https://riverml.xyz/dev/api/cluster/DenStream/) or [CluStream - River](https://riverml.xyz/dev/api/cluster/CluStream/). I haven’t worked with those packages yet, but their features look promising.

---

<div class="post-metadata">

### Author: ![vvbond](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vvbond/32/10105_2.png) [@vvbond](https://discourse.julialang.org/u/vvbond)
#### Post date: [May 1, 2024, 10:28am UTC](https://discourse.julialang.org/t/looking-for-julia-implementations-of-online-clustering-algorithms/113649/5 "2024-05-01T10:28:41Z")

</div>

Thanks for the reference!  
The need for online clustering is not so the data size per se, but a low latency requirement of the real time system I’m developing.
