# DataFrames and MLJ

**URL:** <https://discourse.julialang.org/t/dataframes-and-mlj/122847>\
**Category:** General Usage\
**Tags:** distributed, dataframes, mlj\
**Created:** [November 20, 2024, 11:51am UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847 "2024-11-20T11:51:38Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jose\_Ghislain\_Quenum](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jose_ghislain_quenum/32/204363_2.png) [@Jose\_Ghislain\_Quenum](https://discourse.julialang.org/u/Jose_Ghislain_Quenum)\
**Post date:** [November 20, 2024, 11:51am UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847/1 "2024-11-20T11:51:38Z")

</div>

Hi all,  
I am experiencing a weird behavior from Pluto. I have a dataset that contains more than 1M observations, although the size is about 75M. I’ve been able to cleanup the dataset and address the missing values. However, when I use the ContinuousEncoder to (1) to one-hot encode the categorical values and ensure the continuous values are enforced, it bloats my dataset to a point where a PCA for dimensionality reduction won’t work. I get a message on Pluto that the process has exited. It seems that the underlying Malt worker has crashed.  
What’s the right way to execute this involving the Distributed package? what’s the correct wrokflow to process data frames in MLJ using distribted or multi-threaded processes? Feel free to share some pointers or examples.  
Thanks

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [November 20, 2024, 1:05pm UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847/2 "2024-11-20T13:05:16Z")

</div>

That doesn’t sound like a Pluto issue, but simply OOM? Not sure what you mean by “1 M observations, although the size s about 75M” but if you’ve got 75m rows and you one-hot encode some vector, generating a bunch of additional columns, it’s maybe not surprising to run out of memory?

In any event probably useful to just run your code in the REPL to try and get an error message that isn’t obscured by Malt.

---

<div class="post-metadata">

**Author:** ![Jose\_Ghislain\_Quenum](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jose_ghislain_quenum/32/204363_2.png) [@Jose\_Ghislain\_Quenum](https://discourse.julialang.org/u/Jose_Ghislain_Quenum)\
**Post date:** [November 24, 2024, 8:42am UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847/3 "2024-11-24T08:42:56Z")

</div>

Sorry! The way I put it is misleading. I actually meant to indicate that I was running the code from Pluto.  
The actual issue is with the one-hot encoding, which bloats the dataset and then I suspect the memory allocation fails. Thus, the process exits…  
What I really wanted from the post was to find out what workflow you guys use, especially a distributed version, when handling a large dataset with DataFrames and MLJ  
Regards

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [November 24, 2024, 7:01pm UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847/4 "2024-11-24T19:01:02Z")

</div>

One-hot code bloating is exasperated by any high cardinality features you have. You could try entity embedding: instead of a fixed high-dimensional representation of a feature class, you _learn_ a lower dimensional representation by training using a supervised neural network (NN) that includes embedding layers. The NN may not give the best predictive performance, but once you have the embeddings, you can use these instead of one-not-encoding with whatever supervised model you like. (EvoTreesClassifier or EvoTreesRegressor, the Julia native gradient tree boosters, are pretty good first choices for structured data.).

The NN models provided by MLJFlux now provide entity embedding; an example and citation of the Entity Embedding paper is [here](https://github.com/FluxML/MLJFlux.jl/pull/278/files#diff-9999a08e565eefd64b8b83163587635914c27d4ad3b9aef7278b36d0dc0b1287). IIdeally, you should think about the dimension you need for each multi-class variable, and specify these explicitly, as the defaults are just crude caps on the dimension.

@EssamWisam is working on a version of these models that can be used in a pipeline, but for now you will need to separately train and apply transform using the NN before passing this on manually to your supervised model of choice.

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [November 24, 2024, 7:04pm UTC](https://discourse.julialang.org/t/dataframes-and-mlj/122847/5 "2024-11-24T19:04:31Z")

</div>

To iron out any issues, I suggest you try this out first with a much smaller version of your dataset on a single process.
