# MLJ - A machine learning toolbox for Julia

**URL:** https://discourse.julialang.org/t/mlj-a-machine-learning-toolbox-for-julia/23681
**Category:** Package Announcements
**Created:** [April 30, 2019, 3:25am UTC](https://discourse.julialang.org/t/mlj-a-machine-learning-toolbox-for-julia/23681 "2019-04-30T03:25:12Z")
**Posts on this page:** 1
**Page:** 1

<div class="post-metadata">

### Author: ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)
#### Post date: [April 30, 2019, 3:25am UTC](https://discourse.julialang.org/t/mlj-a-machine-learning-toolbox-for-julia/23681/1 "2019-04-30T03:25:12Z")

</div>

# MLJ - Machine Learning in Julia

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/9/e977632350ed7464efc76c25a4fa38fd0493fbfc.png)

[MLJ](https://github.com/alan-turing-institute/MLJ.jl) is a new flexible framework for composing and tuning supervised and unsupervised learning models, currently scattered in assorted Julia packages, as well as wrapped models from other languages. The MLJ project also seeks to focus efforts in the Julia ML community, and in particular to help inter-operability and maintainability of key ML packages.

The package has been developed primarily at [The Alan Turing Institute](https://turing.ac.uk) but enjoys a growing list of advisors and [contributors](https://github.com/alan-turing-institute/MLJ.jl/blob/master/CONTRIBUTE.md). If you like the project, please star the GitHub [repo](https://github.com/alan-turing-institute/MLJ.jl) to boost the prospects of a **pending funding review.**

### Quick links

☞ [MLJ vs ScikitLearn.jl](https://alan-turing-institute.github.io/MLJ.jl/dev/frequently_asked_questions/)

☞ Video from [London Julia User Group meetup in March 2019](https://www.youtube.com/watch?v=CfHkjNmj1eE) (skip to [demo at 21’39](https://youtu.be/CfHkjNmj1eE?t=21m39s)) &nbsp;

☞ [Basic Usage](https://alan-turing-institute.github.io/MLJ.jl/dev/) and [Tour](https://github.com/alan-turing-institute/MLJ.jl/blob/master/docs/src/tour.ipynb)

☞ Building a [self-tuning random forest](https://github.com/alan-turing-institute/MLJ.jl/blob/master/examples/random_forest.ipynb)

☞ An MLJ [docker image](https://github.com/ysimillides/mlj-docker) (including tour)

☞ Implementing the MLJ interface for a [new model](https://alan-turing-institute.github.io/MLJ.jl/dev/adding_models_for_general_use/)

☞ How to [contribute](https://github.com/alan-turing-institute/MLJ.jl/blob/master/CONTRIBUTE.md)

☞ Julia [Slack](http://julialang.slack.com) channel: #mlj.

### Key implemented features

- **Learning networks.** Flexible model composition beyond traditional  
pipelines (more on this below).

- **Automatic tuning.** Automated tuning of hyperparameters, including  
composite models. Tuning implemented as a model wrapper for  
composition with other meta-algorithms.

- **Homogeneous model ensembling.**

- **Registry for model metadata.** Metadata available without loading  
model code. Basis of a “task” interface and facilitates  
model composition.

- **Task interface.** Automatically match models to specified learning  
tasks, to streamline benchmarking and model selection.

- **Clean probabilistic API.** Improves support for Bayesian  
statistics and probabilistic graphical models.

- **Data container agnostic.** Present and manipulate data in your  
favorite Tables.jl format.

- **Universal adoption of categorical data types.** Enables model  
implementations to properly account for classes seen in training but  
not in evaluation.

### Some planned enhancements

- Integrate **deep learning** packages, such as Flux.jl.

- Model agnostic gradient descent tuning using **automatic  
differentiation**.

- Enhance support for time series and sparse data.

- Add support for heterogeneous/distributed architectures.

- Package common learning network architectures ( **linear pipelines** ,  
**stacks** , etc) as simple one-line operations.

- Implement systematic **benchmarking** for models matching a given task.

- Automated estimates of **cpu and memory requirements** for given task/model.

- Implement DAG style **scheduling**.

- Extend and integrate existing **loss function** libraries to better handle  
probabilistic prediction.

- Add **interpretable machine learning** measures.

- Add **online learning** support.

Feedback, and offers of help very welcome!
