# MLJ/MljFlux standardisation of variables in each cross-validation fold

**URL:** <https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093>\
**Category:** New to Julia\
**Tags:** question, mlj, standardized, mljflux\
**Created:** [March 8, 2026, 3:24pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093 "2026-03-08T15:24:33Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pablo\_G-D](https://avatars.discourse-cdn.com/v4/letter/p/9fc29f/32.png) [@Pablo\_G-D](https://discourse.julialang.org/u/Pablo_G-D)\
**Post date:** [March 8, 2026, 3:24pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093/1 "2026-03-08T15:24:33Z")

</div>

Hi all,

I am learning Julia, and so far I have implemented a few Bayesian models using Turing and other packages. I found it quite fast and reliable.

I am now building a neural network model for a land-use/land-cover task, and I want to ensure that my six predictive variables are standardised within each stratified fold to avoid data leakage. In other words, I want to standardise the dataset independently within each stratified fold and apply the scaling scores from the training dataset to the holdout dataset in that fold. I have used the ‘pipeline’ construct, but I am unsure if it is doing what I want:

```julia-auto
##### Independent covariates
##### Select columns with raw values of 6 covariates

ind_covs = DataFrames.select(lulc, 4, 5, 6, 7, 8, 9)

ind_covs = DataFrames.rename(ind_covs, [:ndvi, :tree_cover, :ntree_cover, :elev, :slope, :pop])

nrow(ind_covs)

print(describe(ind_covs, :all))

##### Neural network classifier
NeuralNetworkClassifier = @load NeuralNetworkClassifier pkg = MLJFlux

##### Define the neural network classifier
lulc_classifier = NeuralNetworkClassifier(builder = MLJFlux.MLP(; hidden=(6,6), σ=Flux.relu),
            epochs=100, loss = Flux.Losses.crossentropy)

##### Standardise (- is this correct?)
##### Is this doing what I want it to do? 
stdz_classifier = Pipeline(Standardizer(), lulc_classifier)

##### Neural net
nnet_lulc = machine(stdz_classifier, ind_covs, land_use)

##### Evaluate the predictive perfomance of the nnet
mod_perf = evaluate!(nnet_lulc, resampling = StratifiedCV(; nfolds=10), 
                               measure=[balanced_accuracy, cross_entropy])

```

Any tips and suggestions will be more than welcome.

Thanks!

Pablo

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 8, 2026, 7:28pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093/2 "2026-03-08T19:28:52Z")

</div>

Yes, this does what you think it does: the standardisation and superivised learning use only the data from each training fold.

BTW, if you were doing regression, you could additionally standardise the target, using the `TransformedTargetModel` wrapper ([docs](https://juliaai.github.io/MLJ.jl/dev/target_transformations/#Target-Transformations)).

---

<div class="post-metadata">

**Author:** ![Pablo\_G-D](https://avatars.discourse-cdn.com/v4/letter/p/9fc29f/32.png) [@Pablo\_G-D](https://discourse.julialang.org/u/Pablo_G-D)\
**Post date:** [March 9, 2026, 3:55pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093/3 "2026-03-09T15:55:23Z")

</div>

Thanks! Just to double-check that I understand it correctly – is the procedure standardising the test dataset too? If that is the case, there is no data leakage between the training and testing datasets, right?

Thank you very much

Pablo

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 9, 2026, 6:40pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093/4 "2026-03-09T18:40:45Z")

</div>

Yes, for each _test_ fold, the normalisation applied to it is performed using the mean and standard deviation learned from the corresponding _train_ fold. So no data leakage.

---

<div class="post-metadata">

**Author:** ![Pablo\_G-D](https://avatars.discourse-cdn.com/v4/letter/p/9fc29f/32.png) [@Pablo\_G-D](https://discourse.julialang.org/u/Pablo_G-D)\
**Post date:** [March 10, 2026, 3:54pm UTC](https://discourse.julialang.org/t/mlj-mljflux-standardisation-of-variables-in-each-cross-validation-fold/136093/5 "2026-03-10T15:54:57Z")

</div>

Outstanding! Thank you very much
