# Saving models in MLJ - only final ones without data

**URL:** <https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873>\
**Category:** Machine Learning\
**Tags:** question, mlj\
**Created:** [September 7, 2022, 6:23am UTC](https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873 "2022-09-07T06:23:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![julien\_goo](https://avatars.discourse-cdn.com/v4/letter/j/ce73a5/32.png) [@julien\_goo](https://discourse.julialang.org/u/julien_goo)\
**Post date:** [September 7, 2022, 6:23am UTC](https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873/1 "2022-09-07T06:23:56Z")

</div>

Hi all,

I have built ML pipelines using MLJ to predict on live data (randomforest models for now + time serie crossval + grid search of hyperparams). It works like a charm.

However, I have a question. I have to serialize tens of individual MLJ.Machines and also to load them in RAM to serve predictions fast.

To reduce RAM footprint and accelerate the serialization/deserialization, how can I save and load only a Machine with the final model (the trained MLJ pipeline, not to loose the features engineering in the MLJ.pipeline) to apply a MLJ.predict? Without storing data from the train samples or keeping the non-final models (no need to keep all the tested models in the CV). My serialized machines weight around 25Mo for randomforests. I guess I have unnecessary content in them.

Thanks for your advice!  
Ju

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [September 9, 2022, 3:33am UTC](https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873/2 "2022-09-09T03:33:48Z")

</div>

Thanks for giving MLJ a spin.

If you are going to serialise a machine `mach`, the recommended protocol is:

1. define `bare_mach = serializable(mach)`.

2. Use your favourite serialiser on `bare_mach`.

Step 1 ensures that `bare_mach` has all traces of training data removed and ensures the learned parameters have a persistent representation (relevant for the `XGBoost` models, which by default only stores a C pointer).

After deserialising, to obtain `mach_deserialized`, you need to call `restore!` on it before you can reliably call `predict` or whatever on it.

There is a shorter workflow if you are happy to use Julia’s built-in serialiser. Details and examples are [here](https://alan-turing-institute.github.io/MLJ.jl/dev/machines/#Saving-machines).

Does this answer your question?

---

<div class="post-metadata">

**Author:** ![julien\_goo](https://avatars.discourse-cdn.com/v4/letter/j/ce73a5/32.png) [@julien\_goo](https://discourse.julialang.org/u/julien_goo)\
**Post date:** [September 9, 2022, 5:26am UTC](https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873/3 "2022-09-09T05:26:04Z")

</div>

Hi,  
Thanks for the answer. The logic is clear, but it is difficult to get rid of the training data in practice (tracking data traces in nested structures and specific to the chosen model).  
I already use the shorter workflow with:

```julia
# Serialization
MLJ.save("model.jls", mach)
# Deserialization
mach = machine("model.jls")

```

The tricky point is to find and remove nested datasets in the `machine` (I found some datasets in `mach.fitresult.data`, `mach.fitresult.resampled_data` and also in various forms in other fields).  
An option such as `MLJ.save("model.jls, mach, predict_only=true)"` would be super helpful to have a minimal machine for predictions, fast to serialize and to serve predictions behind an API on the fly.

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [September 9, 2022, 5:58am UTC](https://discourse.julialang.org/t/saving-models-in-mlj-only-final-ones-without-data/86873/4 "2022-09-09T05:58:23Z")

</div>

> (tracking data traces in nested structures and specific to the chosen model).

`serializable` is supposed to remove _all_ data, even from nested structures. For example, if the model is a `Composite` then there will be machines associated with an underlying learning network, and any data associated with those machines should also be removed. (The implementation is indeed non-trivial). If you are still seeing data, then please open an issue with minimal working example.
