# Package for clustering data points

**URL:** https://discourse.julialang.org/t/package-for-clustering-data-points/83154
**Category:** Machine Learning
**Tags:** question, clustering
**Created:** [June 21, 2022, 11:08pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154 "2022-06-21T23:08:09Z")
**Posts on this page:** 9
**Page:** 1

<div class="post-metadata">

### Author: ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)
#### Post date: [June 21, 2022, 11:08pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/1 "2022-06-21T23:08:09Z")

</div>

Any package recommendations for clustering data points? I see [clustering.jl](https://github.com/JuliaStats/Clustering.jl) exist, however, I don’t see any recent activity on it. Is it still being maintained? Any other package that might be actively maintained?

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [June 22, 2022, 10:20pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/2 "2022-06-22T22:20:39Z")

</div>

My [BetaML.jl](https://github.com/sylvaticus/BetaML.jl) package provides kmeans, kmedoids (hard clustering) and gmm (soft clustering), and these are available also through the MLJ interface…

---

<div class="post-metadata">

### Author: ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)
#### Post date: [June 22, 2022, 10:57pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/3 "2022-06-22T22:57:25Z")

</div>

Thanks. I was not aware of your package. This should work for my purpose. 🙂 So, I am assuming your package does a lot of ML stuff without focusing too much on computational efficiency. For more focused and tailored algorithms you refer to the alternative packages you list on your GitHub page. Is that right? (Just trying to get an idea of how your package differs from the existing ones).

Also, how do I access it through the MLJ interface? Any reference that I can look into? I have never used it before. 😅

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [June 23, 2022, 4:02am UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/4 "2022-06-23T04:02:01Z")

</div>

Yes that’s correct… then everything is relative.  
I am out of my pc now, but I have some benchmark that eg for missing imputation compared to R mice, time is quite good.  
RF are really slow compared to DecidionTrees.jl, but that because they use an algorithm where they accept almost everything, including unordered and missing data.

For the MLJ interface refer to its documentation… it adds a bunch of concepts to grasp (machine, scientific types,…) but then it allows a conmon api for all ml models…

---

<div class="post-metadata">

### Author: ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)
#### Post date: [June 23, 2022, 10:12am UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/5 "2022-06-23T10:12:07Z")

</div>

You may also like HorseML.jl by @QGMW22 ([[ANN] HorseML.jl v0.4.0](https://discourse.julialang.org/t/ann-horseml-jl-v0-4-0/77181) and [[ANN] HorseML.jlv0.4.1: Many ML algorithms](https://discourse.julialang.org/t/ann-horseml-jlv0-4-1-many-ml-algorithms/80703))

---

<div class="post-metadata">

### Author: ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)
#### Post date: [June 23, 2022, 1:08pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/6 "2022-06-23T13:08:33Z")

</div>

@sylvaticus I have to admit that I just went very briefly though a documentation of BetaML and there seems to be some areas that I have to do additional reading thus I would like to take some of your time and ask you directly if BetaML is automatically choosing the number of clusters?

Also @math_opt, I cannot comment on timings re large data sets but I took a look at my notes and I see that I have been also using ParallelKMeans.jl ([GitHub - PyDataBlog/ParallelKMeans.jl: Parallel & lightning fast implementation of available classic and contemporary variants of the KMeans clustering algorithm](https://github.com/PyDataBlog/ParallelKMeans.jl)). Hope it helps (maybe).

---

<div class="post-metadata">

### Author: ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)
#### Post date: [June 23, 2022, 1:17pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/7 "2022-06-23T13:17:20Z")

</div>

Hi, it doesn’t “automatically” provide the number of clusters, but GMM returns both the Bayesian information criterion (BIC) and the Akaike information criterion (AIC) criteria that can be used to choose it…

---

<div class="post-metadata">

### Author: ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)
#### Post date: [June 23, 2022, 1:40pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/8 "2022-06-23T13:40:25Z")

</div>

Hi, I got it, thanks! (BTW, also many thanks for [Introduction to Scientific Programming and Machine Learning with Julia](https://sylvaticus.github.io/SPMLJ/stable/#Introduction-to-Scientific-Programming-and-Machine-Learning-with-Julia))

---

<div class="post-metadata">

### Author: ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)
#### Post date: [June 23, 2022, 5:22pm UTC](https://discourse.julialang.org/t/package-for-clustering-data-points/83154/9 "2022-06-23T17:22:55Z")

</div>

Thanks for directing me to these packages! I need to look into this. It looks like I am not aware of a lot of existing packages. Glad I asked here.
