# \[ANN\] MLJ: Outlier Detection, Text Analysis, Improved Pipelines and Serialization

**URL:** <https://discourse.julialang.org/t/ann-mlj-outlier-detection-text-analysis-improved-pipelines-and-serialization/79320>\
**Category:** Package Announcements\
**Tags:** machine-learning, mlj, outlier-detection, pipelines, text-analysis\
**Created:** [April 11, 2022, 1:40am UTC](https://discourse.julialang.org/t/ann-mlj-outlier-detection-text-analysis-improved-pipelines-and-serialization/79320 "2022-04-11T01:40:38Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [April 11, 2022, 1:40am UTC](https://discourse.julialang.org/t/ann-mlj-outlier-detection-text-analysis-improved-pipelines-and-serialization/79320/1 "2022-04-11T01:40:38Z")

</div>

There have been quite a few releases (two breaking) since our [last announcement](https://discourse.julialang.org/t/ann-mlj-now-has-iteration-control-and-model-stacking/63726). Here are some highlights, available in MLJ 0.18.

## Tutorial: MLJ for Data Scientists in Two Hours

We have added a [new tutorial](https://github.com/alan-turing-institute/MLJ.jl/blob/dev/examples/telco/notebook.ipynb) focused on the practicing data scientist transitioning from another platform, such as scikit-learn or caret.

## Pipelines without macros

For improved robustness, pipelines are no longer created with a macro. Every pipeline is just an instance of a single parameterized type [`Pipeline`](https://alan-turing-institute.github.io/MLJ.jl/dev/linear_pipelines/), with a constructor of the same name. However, unless you have special options to specify, you just combine models using the `|>` syntax:

```julia-auto
using MLJ

LinearRegressor = @load LinearRegressor pkg=MLJLinearModels add=true
PCA = @load PCA pkg=MultivariateStats add=true
  
pipe1 = MLJBase.table |> ContinuousEncoder |> Standardizer
pipe2 = PCA |> LinearRegressor
pipe3 = pipe1 |> pipe2

```

## Target transformers

Transformation of the target (with post-prediction inversion) is no longer available in pipelines, but provided instead by the [`TransformedTarget(model, ...)`](https://alan-turing-institute.github.io/MLJ.jl/dev/target_transformations/#Target-Transformations) model wrapper.

## Serialization (Olivier Labayle @olivierlabayle)

[Serialization](https://alan-turing-institute.github.io/MLJ.jl/dev/machines/#Saving-machines) in MLJ has changed. The previous `MLJ.save(mach)` method still works, but you can only save using Julia’s native JLS format in this way (and the new format is not backwards compatible). However, a new workflow allows for serialization using any generic serializer; serialization plays nicely with model composition and model wrappers, such as `TunedModel` and `EnsembleModel` (even for non-Julia atomic models); and training data will not be inadvertently serialized.

## Outlier detection (David Muhr @davnn)

MLJ now wraps a large number of outlier detection models from [OutlierDetection.jl](https://outlierdetectionjl.github.io/OutlierDetection.jl/dev/). In MLJ, do `models("Detector")` to list these. See [here](https://outlierdetectionjl.github.io/OutlierDetection.jl/dev/documentation/simple-usage/) for usage.

## Text analysis (Chris Alexander @pazzo83)

MLJ now provides some text analysis tools. See the [MLJText.jl readme](https://github.com/JuliaAI/MLJText.jl#readme) for details. Also new is the `TSVDTransformer` model for truncated singular value decomposition, an interface to [TSVD.jl](https://github.com/JuliaLinearAlgebra/TSVD.jl).
