# Best input/output format for Julia ML packages

**URL:** <https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078>\
**Category:** General Usage\
**Created:** [October 1, 2022, 10:26am UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078 "2022-10-01T10:26:31Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [October 1, 2022, 10:26am UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/1 "2022-10-01T10:26:32Z")

</div>

Following a recent discussion on slack, I’d like to raise this point here on discourse for arguments not to disappear. Also, the slack format with many short messages isn’t the most convenient for longer discussions.

Currently, ML packages tend to require/return opaque matrices of numbers, and even the dimensions order is not consistent. See this example by @jling:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/0/c/0c6d66c71d815d6992b9218c44174f82ac9be0c9.png)

A couple alternatives were discussed, each with pros and cons:

- Arrays with named dimensions, like `:observation` and `:feature` instead of unnamed matrices. Such arrays are well-supported in Julia with a lightweight focused package, eg `NamedDims.jl`.

- Collection-of-collections, such as (abstract)vector of observations, with each observation an (abstract)vector of features.

- Table of objects, with columns representing features.

I’m only a user of ML libraries myself, and find the second alternative above the most natural, and most likely easiest to make efficient. Basically, for ML algos that take a matrix, the adaptation amounts to:

```julia
fit(A::Matrix{<:Real}) = ... original method ...
fit(A::AbstractVector) = fit(stack(A))
# potentially swap these two, depending on the original algo dimension order:
fit(A::ColumnSlices) = fit(parent(A))
fit(A::RowSlices) = fit(permutedims(parent(A)))

```

I wonder what others think about these alternatives, especially ML package developers.

Note that this post focuses on classic ML for now, in DL/NNs where the situation may (?) be somewhat more complicated.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 1, 2022, 1:34pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/2 "2022-10-01T13:34:55Z")

</div>

I think MLJ has a good system for this, standardizing on tables and using ScientificTypes.jl: [https://alan-turing-institute.github.io/MLJ.jl/dev/getting\_started/#Data-containers-and-scientific-types](https://alan-turing-institute.github.io/MLJ.jl/dev/getting_started/#Data-containers-and-scientific-types)

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 1, 2022, 2:54pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/3 "2022-10-01T14:54:13Z")

</div>

except it’s not composable with the data cleaning pipeline other libraries use.

which means you have to do everything in MLJ and some backend has performance problems (with or without GPUs), then you have to re-write the whole data cleaning pipeline in a different way if you were to switch away from MLJ

and for example MLJ doesn’t have XGBoost (it has EvoTree, but EvoTree is not XGBoost and EvoTree can’t do incremental training)

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 1, 2022, 8:33pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/4 "2022-10-01T20:33:18Z")

</div>

It definitely does have XGBoost, there’s MLJXGBoostInterface, which is basically all I’ve been using from MLJ. I haven’t run into composability issues myself, not really sure what you mean.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 1, 2022, 8:40pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/5 "2022-10-01T20:40:46Z")

</div>

> [@ericphanson](#):
>
> composability

data cleaning from Flux-circle doesn’t work with MLJ and vise versa, so if say, MLJFlux has performance problem (it does), you have to rewrite literally everything in non-MLJ again

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 1, 2022, 8:44pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/6 "2022-10-01T20:44:01Z")

</div>

Can’t you do something like this? [code sketch]

```julia
data = ...get data...
my_clean!(data)
data = rearrange_data_so_MLJ_is_happy(data) # maybe needed?
my_model = ...create MLJ machine with `data`...
MLJ.fit!(my_model)

```

Or is this what you mean by

> you have to rewrite literally everything in non-MLJ again

?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 1, 2022, 8:54pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/7 "2022-10-01T20:54:16Z")

</div>

MLJ uses different packages and functions for data splitting and labeling and augmentation; I can show an example later, no promise thought 😃

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [October 1, 2022, 9:35pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/8 "2022-10-01T21:35:10Z")

</div>

> [@ericphanson](#):
>
> I think MLJ has a good system for this, standardizing on tables and using ScientificTypes.jl

Tables probably work fine for simple cases, but they seem more difficult to generalize: varying number of features, multidimensional datasets, online processing…

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 1, 2022, 10:19pm UTC](https://discourse.julialang.org/t/best-input-output-format-for-julia-ml-packages/88078/9 "2022-10-01T22:19:15Z")

</div>

I think that might be true but you can also pick and choose what to use from it and what not to. Eg I am using my own cross validation code even though MLJ has some (because I want to get the predictions back per fold to do some more evaluation on). To me it doesn’t seem like there’s a fundamental mismatch except to the extent you need to convert things to their api and back to whatever form you want to use it.
