# Machine Learning Classification

**URL:** <https://discourse.julialang.org/t/machine-learning-classification/92915>\
**Category:** Machine Learning\
**Tags:** question\
**Created:** [January 13, 2023, 11:24am UTC](https://discourse.julialang.org/t/machine-learning-classification/92915 "2023-01-13T11:24:47Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![BadBoy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/badboy/32/202815_2.png) [@BadBoy](https://discourse.julialang.org/u/BadBoy)\
**Post date:** [January 13, 2023, 11:24am UTC](https://discourse.julialang.org/t/machine-learning-classification/92915/1 "2023-01-13T11:24:47Z")

</div>

How can I find the best subset of the predictors for a classification problem??  
Any kind of help will be appreciated. Thanks in advance.

---

<div class="post-metadata">

**Author:** ![LucasMSpereira](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lucasmspereira/32/26393_2.png) [@LucasMSpereira](https://discourse.julialang.org/u/LucasMSpereira)\
**Post date:** [January 13, 2023, 1:05pm UTC](https://discourse.julialang.org/t/machine-learning-classification/92915/2 "2023-01-13T13:05:45Z")

</div>

If I understood correctly, you should look for unsupervised learning, explainable AI and dimensionality reduction techniques. Possible starting points are: [MLJ.jl models](https://alan-turing-institute.github.io/MLJ.jl/dev/list_of_supported_models/), [ExplainableAI.jl](https://github.com/adrhill/ExplainableAI.jl), [ShapML.jl](https://github.com/nredell/ShapML.jl) and [CounterfactualExplanations.jl](https://github.com/pat-alt/CounterfactualExplanations.jl)

---

<div class="post-metadata">

**Author:** ![BadBoy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/badboy/32/202815_2.png) [@BadBoy](https://discourse.julialang.org/u/BadBoy)\
**Post date:** [January 13, 2023, 1:23pm UTC](https://discourse.julialang.org/t/machine-learning-classification/92915/3 "2023-01-13T13:23:37Z")

</div>

No, no. When we are build a model to dataset, it is not the case that all the predictors that are given in the data is important. So, to choose the best subset of the set of predictors, i have learnt four techniques best subset selection, forward selection, backward selection and hybrid selection. So, how we can apply these methods, that was my question. No need to go for unsupervised learning.

---

<div class="post-metadata">

**Author:** ![LucasMSpereira](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lucasmspereira/32/26393_2.png) [@LucasMSpereira](https://discourse.julialang.org/u/LucasMSpereira)\
**Post date:** [January 13, 2023, 1:41pm UTC](https://discourse.julialang.org/t/machine-learning-classification/92915/4 "2023-01-13T13:41:27Z")

</div>

Oh ok. That’s [feature selection](https://en.wikipedia.org/wiki/Feature_selection). A quick google search resulted in [FeatureSelectors.jl](https://github.com/darrencl/FeatureSelectors.jl). Scikit-learn has nice [options](https://scikit-learn.org/stable/modules/feature_selection.html) (julia package [ScikitLearn.jl](https://cstjean.github.io/ScikitLearn.jl/dev/man/pipelines/#) may enable access to them)
