# \[ANN\] MLJ: an update

**URL:** <https://discourse.julialang.org/t/ann-mlj-an-update/31676>\
**Category:** Machine Learning\
**Created:** [November 29, 2019, 10:07pm UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676 "2019-11-29T22:07:32Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [November 29, 2019, 10:07pm UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/1 "2019-11-29T22:07:32Z")

</div>

A **lot** of work has been done on [MLJ](https://github.com/alan-turing-institute/MLJ.jl) and [associated repos](https://github.com/alan-turing-institute/MLJ.jl#the-mlj-universe) since @ablaom’s [last announcement](https://discourse.julialang.org/t/mlj-a-machine-learning-toolbox-for-julia/23681) some 7 months ago so here’s a brief update:

- MLJ now interfaces with around [100 models](https://github.com/alan-turing-institute/MLJ.jl/blob/master/README.md#models-available) including most of ScikitLearn’s,
- the `@pipeline` macro makes it easy for users to define simple sequence of operations and models,
- [extensive documentation](https://alan-turing-institute.github.io/MLJ.jl/stable/) and a new dedicated repository for tutorials: [MLJTutorials](https://github.com/alan-turing-institute/MLJTutorials), which includes end-to-end examples and a port of the _Introduction to Statistical Learning_’s labs,
- [MLJBase](https://github.com/alan-turing-institute/MLJBase.jl) now supports many useful metrics for regression and classification; it also makes it easier for package devs to work with multiple tabular data formats (via [Tables.jl](https://github.com/JuliaData/Tables.jl)) and categorical features (via [CategoricalArrays.jl](https://github.com/JuliaData/CategoricalArrays.jl)).

### For devs

MLJ can help you focus on “just” developing models and benefit from MLJ’s machinery for data pre-processing, hyper-parameter tuning, evaluation metrics, etc.

If you have or know of a Julia package that fits the idea of the `fit/predict/transform`, please consider adding an interface to MLJ and registering your package with [MLJModels](https://github.com/alan-turing-institute/MLJModels.jl). This will allow users to discover, compare and compose (many) models.

Thanks a lot to those who’ve already been helping us out and/or have given us detailed feedback on MLJ (@samuel_okon, @nignatiadis, @juliohm, @cscherrer, @ZacLN, @jpsamaroo, and many others)

### Brief roadmap

In future months we will be focusing on

- polishing the user interface: we would like to stabilise the API for MLJ and MLJBase by the end of February and release a 1.0 for both then,
- adding capacity for more sophisticated hyper parameter tuning,
- improving MLJ’s support for distributed & multithreaded computing.

See also the [suggested projects](https://github.com/alan-turing-institute/MLJ.jl#contributing-to-mlj) section for contributors if you’re interested in helping out (thanks!).

Have a nice weekend

---

<div class="post-metadata">

**Author:** ![robsmith11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robsmith11/32/29641_2.png) [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Post date:** [November 30, 2019, 7:55am UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/2 "2019-11-30T07:55:41Z")

</div>

You mentioned “sophisticated hyper parameter tuning” in the road map. What methods are you considering and/or already supporting?

Anything along the line of R’s [DiceOptim](https://cran.r-project.org/web/packages/DiceOptim/index.html) for optimization with noise?

---

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [November 30, 2019, 9:54am UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/3 "2019-11-30T09:54:08Z")

</div>

Well anything more sophisticated than grid or basic random search (which we support). For instance we would like to try using/interfacing with [Hyperopt.jl](https://github.com/baggepinnen/Hyperopt.jl) (which supports BO, fancier random search, etc). I’m not familiar with R’s DiceOptim (thanks for the link).  
It would also be nice is to try to leverage appropriate tools from JuliaOpt to help for such things too.

---

<div class="post-metadata">

**Author:** ![Ratingulate](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ratingulate/32/9242_2.png) [@Ratingulate](https://discourse.julialang.org/u/Ratingulate)\
**Post date:** [December 1, 2019, 3:55am UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/4 "2019-12-01T03:55:13Z")

</div>

Exciting release!

Any thoughts on streaming and distributed/out of core datasets?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 1, 2019, 4:04am UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/5 "2019-12-01T04:04:15Z")

</div>

It would be great to have general support in MJL.jl for such things. I have made a prototype in JLBoost.jl where the user can use `JDF.JDFFile` to fit gradient-boosting tree (same family as XGBoost) on larger-than-RAM datasets in an out-of-core fashion. See [https://github.com/xiaodaigh/JLBoost.jl#fit-model-on-jdfjdffile---enabling-larger-than-ram-model-fit](https://github.com/xiaodaigh/JLBoost.jl#fit-model-on-jdfjdffile---enabling-larger-than-ram-model-fit)

But it’s not MLJ.jl though. Keen to bring such support to my packages once MLJ has a general framework for it.

---

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [December 1, 2019, 11:46am UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/6 "2019-12-01T11:46:10Z")

</div>

> [@Ratingulate](#):
>
> Any thoughts on streaming and distributed/out of core datasets?

On the streaming side, interop with OnlineStats is considered (maybe @ablaom can comment on this).

With respect to distributed/out-of-core, MLJ is built around compatibility with the Tables.jl interface; it works with JuliaDB for instance, and it would work with anything else that corresponds to Tables.jl

---

<div class="post-metadata">

**Author:** ![datnamer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datnamer/32/3471_2.png) [@datnamer](https://discourse.julialang.org/u/datnamer)\
**Post date:** [December 1, 2019, 2:08pm UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/7 "2019-12-01T14:08:53Z")

</div>

What happens to the data after it leaves the table? Is mlj typed for abstract arrays or dense in memory?

---

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [December 1, 2019, 10:55pm UTC](https://discourse.julialang.org/t/ann-mlj-an-update/31676/8 "2019-12-01T22:55:31Z")

</div>

Maybe it’s useful to stress that MLJ is effectively “just” a tool to use and compose _other_ models, some are offered by MLJ but most are external. In the interface to a given package, MLJ expects a Table object by design but then the implementation of the `fit` function is free to do whatever it wants so that the data is in a form that can be processed by the model.  
To facilitate this, we offer functions that facilitate conversion of the data _if required_ by the package and, for instance, to matrices with or without transposition depending on what’s required.

To exemplify this, a number of algorithms in the Stats environment expect a `p x n` _matrix_ as input convention; the MLJ interface can provide this as an adjoint view (the convention in MLJ is `n x p`); in the unfortunate case where the algorithm does not support adjoints, then a copy is needed but, fingers crossed, this will slowly be updated in MultivariateStats and the like so that we don’t have to copy data around too much if not required.

I hope this makes sense and answers your question
