# Right way of applying \`inverse\_transform\`

**URL:** <https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124>\
**Category:** Machine Learning\
**Tags:** mlj, standardized\
**Created:** [June 21, 2022, 1:19pm UTC](https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124 "2022-06-21T13:19:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ctrebbau](https://avatars.discourse-cdn.com/v4/letter/c/5f9b8f/32.png) [@ctrebbau](https://discourse.julialang.org/u/ctrebbau)\
**Post date:** [June 21, 2022, 1:19pm UTC](https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124/1 "2022-06-21T13:19:43Z")

</div>

Hi, I have a setup like this:

```julia
dftrain, dftest = partition(df, 0.7, shuffle=true, rng=123)

datapipe = ContinuousEncoder() |> Standardizer()
datatrans_mach = machine(datapipe, dftrain) |> fit!
normalized_train = MLJ.transform(datatrans_mach, dftrain)
normalized_test = MLJ.transform(datatrans_mach, dftest)

normalizer = fitted_params(datatrans_mach).machines[2]

ytrain, Xtrain = normalized_train.target, select(normalized_train, Not(target))
ytest, Xtest = normalized_test.target, select(normalized_test, Not(target))

knn = KNNRegressor()
knnM = machine(knn, Xtrain, ytrain) |> fit!

```

All well and good, but when I do

```julia
predict(knn, inverse_transform(normalizer, Xtest))

```

I get  
`ERROR: Attempting to transform data with incompatible feature labels.`  
So I thought that since `Standardizer` was trained on `dftrain` which includes the `target` I tried

```julia
predict(knn, inverse_transform(normalizer, hcat(Xtest, ytest)))

```

But that is evidently not it, since I get a more fundamental incompatibility `ERROR: ArgumentError: dimension of input points:44 and tree data:43 must agree`

And if I first predict and then inverse transform i.e.

```julia
inverse_transform(normalizer, predict(knn, Xtest))

```

I get `ERROR: type Nothing has no field names`

So, how can I get the predictions on the original scale?

---

<div class="post-metadata">

**Author:** ![svilupp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/svilupp/32/34933_2.png) [@svilupp](https://discourse.julialang.org/u/svilupp)\
**Post date:** [June 22, 2022, 12:51pm UTC](https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124/2 "2022-06-22T12:51:48Z")

</div>

> [@ctrebbau](#):
>
> All well and good, but when I do
> 
> ```julia
> predict(knn, inverse_transform(normalizer, Xtest))
> 
> ```

What are you hoping to do?

If I understood your example, Xtrain is your normalized training dataset and Xtest is your normalized testing dataset.  
So if you trained your model on Xtrain, you should be able to do predict() on Xtest like this

```julia
# get predictions for your test dataset
ytest_hat=predict(knnM, Xtest)

```

* * *

It is possible that your error is simply a typo, because your fitted machine is called `knnM` whereas your predict() call is against `knn`

* * *

In general, would it be possible for you to change your workflow and separate your X and y early on (ala [Common MLJ Workflows](https://alan-turing-institute.github.io/MLJ.jl/dev/common_mlj_workflows/#Data-ingestion) )

That way your target transformations would be separate and would be easy to debug if you have any problems. I’d argue it’s the more common way, because there are transforms that could introduce leakage from target into your features, so you tend to separate those early on.

Eg, changing your code to:

```julia
y, X = unpack(df, ==(:target), rng=123);
(Xtrain, Xtest), (ytrain, ytest) = partition((X, y), 0.7, shuffle=true,multi=true, rng=123)

datapipe = ContinuousEncoder() |> Standardizer()
datatrans_mach = machine(datapipe, Xtrain) |> fit!

normalized_train = MLJ.transform(datatrans_mach, Xtrain)
normalized_test = MLJ.transform(datatrans_mach, Xtest);

knn = KNNRegressor()
knnM = machine(knn, normalized_train, ytrain) |> fit!

# out of sample predictions that you can evaluate performance on
ytest_hat=predict(knnM, normalized_test)

```

---

<div class="post-metadata">

**Author:** ![samuel\_okon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/samuel_okon/32/10285_2.png) [@samuel\_okon](https://discourse.julialang.org/u/samuel_okon)\
**Post date:** [June 23, 2022, 3:36pm UTC](https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124/3 "2022-06-23T15:36:44Z")

</div>

@ctrebbau I’m not sure I get your question. But check to see if the following code does what you want. You could use whatever workflow you wish to but using the MLJ workflows @svilupp as pointed out makes things easier conceptually.

```julia
# This assumes that the name of your target feature is `:target`
# You can replace this with the actual name of your target feature
knnp = (X -> select(X, Not(:target))) |> KNNRegressor
knnM = machine(knnp, normalized_train, ytrain) |> fit!
predict(knnM, inverse_transform(normalizer, normalized_test))

```

---

<div class="post-metadata">

**Author:** ![ctrebbau](https://avatars.discourse-cdn.com/v4/letter/c/5f9b8f/32.png) [@ctrebbau](https://discourse.julialang.org/u/ctrebbau)\
**Post date:** [June 25, 2022, 6:27pm UTC](https://discourse.julialang.org/t/right-way-of-applying-inverse-transform/83124/4 "2022-06-25T18:27:18Z")

</div>

Hi, thank you for your prompt help; sorry for my late response. I’m sorry I wasn’t able to explain myself more clearly. I’ve adhered more closely to the more standard workflow, separating target and features earlier, even before normalizing, and I’m happy to report I’m getting more sensible predictions now.
