# Pipeline for different columns

**URL:** <https://discourse.julialang.org/t/pipeline-for-different-columns/81494>\
**Category:** Machine Learning\
**Created:** [May 23, 2022, 7:52am UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494 "2022-05-23T07:52:41Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 23, 2022, 7:52am UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/1 "2022-05-23T07:52:41Z")

</div>

Hi everyone, I have mixed type data and want to apply different functions to different types in a pipeline.  
Like imputation with mean for continuous and knn imputation for categorical and so on…  
Is it possible with MLJ or some other package?

Thanks

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [May 23, 2022, 10:51am UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/2 "2022-05-23T10:51:56Z")

</div>

Are you aware of

> **[GitHub - JuliaML/TableTransforms.jl: Transforms and pipelines with tabular...](https://github.com/JuliaML/TableTransforms.jl)**
>
> Transforms and pipelines with tabular data in Julia - GitHub - JuliaML/TableTransforms.jl: Transforms and pipelines with tabular data in Julia

We are already using ScientificTypes.jl to decide whether or not it makes sense to apply certain transforms to continuous vs. categorical columns. We could easily provide a generic transform that knows which columns to select, but this is use-case-specific. Feel free to reach us out in our Zulip #machine-learning stream for further questions.

---

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 23, 2022, 1:21pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/3 "2022-05-23T13:21:32Z")

</div>

Thanks for the reply.

From what I am seeing, it applies automatically from the scientific types right? There is no way for manually selecting?

Thank for the input, i will check the package out.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [May 23, 2022, 1:24pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/4 "2022-05-23T13:24:01Z")

</div>

In DataFrames.jl, you can select columns by type, i.e.

```julia
julia> using DataFrames

julia> df = DataFrame(a = [1, 2], b = ["x", "y"], c = [1.5, 3.5]);

julia> df[:, names(df, String)]
2×1 DataFrame
 Row │ b      
     │ String 
─────┼────────
   1 │ x
   2 │ y

julia> df[:, names(df, Real)]
2×2 DataFrame
 Row │ a c       
     │ Int64 Float64 
─────┼────────────────
   1 │ 1 1.5
   2 │ 2 3.5

```

So you can subset columns as needed and operate on them. You can use the same strategy inside `transform` calls.

---

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 23, 2022, 1:34pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/5 "2022-05-23T13:34:07Z")

</div>

Thanks for the reply but I was aiming for a method that could include inside a pipeline in order to create a more straightfoward and compiled way of treating data.  
more like this:  
[https://alan-turing-institute.github.io/MLJ.jl/dev/linear\_pipelines/](https://alan-turing-institute.github.io/MLJ.jl/dev/linear_pipelines/)  
and the columns transformers from scikit-learn in python

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [May 23, 2022, 3:31pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/6 "2022-05-23T15:31:58Z")

</div>

> [@jfcal](#):
>
> From what I am seeing, it applies automatically from the scientific types right? There is no way for manually selecting?

You can select/reject columns with `Select` and `Reject` and then apply specific transforms. Later you can join the results with a `Parallel` transforms. The documentation explains this better.

---

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 23, 2022, 4:09pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/7 "2022-05-23T16:09:00Z")

</div>

ok thanks!

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [May 23, 2022, 9:48pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/8 "2022-05-23T21:48:22Z")

</div>

You can do some of what you want in MLJ. The built-in transformers are documented [here](https://alan-turing-institute.github.io/MLJ.jl/dev/transformers/). The `FillImputer` handles mixed types but is pretty basic. There is also `MissingImputator` (an MLJ model with core algorithm provided by BetaML.jl) for continuous data, which uses EM clustering. Probably not good for larger datasets. You might need to look at BetaML.jl to get detailed docs.

Unlike TableTransforms.jl, feature selection is specified as a transformer hyper-parameter. Currently MLJ has not adopted the split/apply/combine paradigm of TableTransforms.jl (although I think that’s probably a good idea).

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [May 23, 2022, 10:47pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/9 "2022-05-23T22:47:40Z")

</div>

And the transformer `OneHotEncoder` now supports `missing` values.

---

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 24, 2022, 7:23am UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/10 "2022-05-24T07:23:21Z")

</div>

Thanks for the reply, @ablaom .

From what i understood from the docs and your reply. MLJ has built-in transformers that handle mixed data types but there is no way of applying certain functions to a certain data (or scientific) type through the pipeline (like tabletransforms.jl) ?

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [May 25, 2022, 9:02pm UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/11 "2022-05-25T21:02:13Z")

</div>

> [@jfcal](#):
>
> but there is no way of applying certain functions to a certain data (or scientific) type through the pipeline

Well, rather than split/apply/combine, you send your table to a model, that model selectively operates on certain columns - for example `OneHotEncoder` just spawns new columns for the `Multiclass` types - and the other columns are left untouched. The next model might selectively standardise all `Continuous` features, and so. So, in principle, you should be able to carry out the same kinds of processing. But you are forced to compute in sequence. So performance may not be as good. I think the TableTransforms.jl approach is better.

That said, if you build a composite model using MLJ’s [learning network](https://alan-turing-institute.github.io/MLJ.jl/dev/composing_models/#Learning-Networks) syntax (instead of the “canned” linear pipeline syntax) then you have more flexibility. You can do split/apply/combine and a lot more. (For example MLJ’s model `Stack` functionality is implemented using learning networks. And there a PR under review to make learning networks multithreading.) But for routine pre-processing, this might be overkill.

---

<div class="post-metadata">

**Author:** ![jfcal](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jfcal](https://discourse.julialang.org/u/jfcal)\
**Post date:** [May 29, 2022, 9:08am UTC](https://discourse.julialang.org/t/pipeline-for-different-columns/81494/12 "2022-05-29T09:08:23Z")

</div>

thanks for the reply. Will try it out!
