# Simple tool for train test split

**URL:** <https://discourse.julialang.org/t/simple-tool-for-train-test-split/473>\
**Category:** Machine Learning\
**Created:** [November 21, 2016, 3:16pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473 "2016-11-21T15:16:15Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![klezmen](https://avatars.discourse-cdn.com/v4/letter/k/a698b9/32.png) [@klezmen](https://discourse.julialang.org/u/klezmen)\
**Post date:** [November 21, 2016, 3:16pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/1 "2016-11-21T15:16:15Z")

</div>

Does anyone know a tool that enables me to split a dataframe in to a training set and a test set?

---

<div class="post-metadata">

**Author:** ![tbreloff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbreloff/32/68_2.png) [@tbreloff](https://discourse.julialang.org/u/tbreloff)\
**Post date:** [November 21, 2016, 4:18pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/2 "2016-11-21T16:18:56Z")

</div>

The future of iteration and dataset partitioning is soon to be merged into MLDataUtils: [Data Access Pattern in 0.5 by Evizero · Pull Request #16 · JuliaML/MLDataUtils.jl · GitHub](https://github.com/JuliaML/MLDataUtils.jl/pull/16).

There isn’t native support for DataFrames but you would only need to define `nobs`/`getobs` for an AbstractDataFrame (or AbstractTable?) and all this functionality will be available.

---

<div class="post-metadata">

**Author:** ![Evizero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/evizero/32/10118_2.png) [@Evizero](https://discourse.julialang.org/u/Evizero)\
**Post date:** [November 21, 2016, 4:29pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/3 "2016-11-21T16:29:43Z")

</div>

I secretly wish this question would not come up for at least a week until I found the time to write the docs and tutorials 🙂

There will be DataFrame support, or at least a tutorial that shows how to use DataFrames, but I need some time to finish this up.

Until then I would advice to simply do it manually. I assume each row represents an observation? If so you could use this little snipped:

```julia
function partitionTrainTest(data, at = 0.7)
    n = nrow(data)
    idx = shuffle(1:n)
    train_idx = view(idx, 1:floor(Int, at*n))
    test_idx = view(idx, (floor(Int, at*n)+1):n)
    data[train_idx,:], data[test_idx,:]
end

using RDatasets
iris = dataset("datasets", "iris")
train,test = partitionTrainTest(iris, 0.7) # 70% train

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [November 21, 2016, 4:32pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/4 "2016-11-21T16:32:59Z")

</div>

Given a dataframe `df`, to pull out e.g. 5% as a training set you could presumably just do:

```julia
sample = randsubseq(1:size(df,1), 0.05)
trainingset = df[sample, :]
notsample = [i for i in 1:size(df,1) if isempty(searchsorted(sample, i))]
testset = df[notsample, :]

```

i.e. using the built-in `randsubseq` function to efficiently get a random (ordered) subset `sample` of the rows, without repetition, where each row is included with probability `0.05`.

---

<div class="post-metadata">

**Author:** ![klezmen](https://avatars.discourse-cdn.com/v4/letter/k/a698b9/32.png) [@klezmen](https://discourse.julialang.org/u/klezmen)\
**Post date:** [November 21, 2016, 10:45pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/5 "2016-11-21T22:45:53Z")

</div>

Thank’s so far, at least i have a solution now.

---

<div class="post-metadata">

**Author:** ![microgold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/microgold/32/5404_2.png) [@microgold](https://discourse.julialang.org/u/microgold)\
**Post date:** [September 16, 2018, 5:03pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/6 "2018-09-16T17:03:05Z")

</div>

I used the MLDataUtils package and finally figured out how to do this with dataframes into a Random Forest:

Here is my sample from my notebook:

```julia
using MLDataUtils

#convert the dataframes into arrays
x = convert(Array,iris[1:4])
y = Array{Int64}(iris[:SpeciesEnumerator])

# shuffle the data so its not in order when we split it up
Xs, Ys = shuffleobs((transpose(x), y))

#now split the data into training sets and validation sets
(X_train1, y_train1), (X_test1, y_test1) = splitobs((Xs, Ys); at = 0.67)
    
 # need to convert the split data back into arrays

    x_train = Array(transpose(X_train1))
    y_train = Array(y_train1)
    x_test = Array(transpose(X_test1))
    y_test = Array(y_test1)

 # Create a random forest model
model = RandomForestClassifier(n_subfeatures = 3, n_trees = 50, partial_sampling=0.7, max_depth = 4)

# Train the model on the dataset 
DecisionTree.fit!(model, x_train, y_train)

# Apply the trained model to the test features data set 
prediction = convert(Array{Int64,1}, DecisionTree.predict(model, x_test))

```

---

<div class="post-metadata">

**Author:** ![Ajaychat3](https://avatars.discourse-cdn.com/v4/letter/a/ecd19e/32.png) [@Ajaychat3](https://discourse.julialang.org/u/Ajaychat3)\
**Post date:** [September 18, 2018, 1:20pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/7 "2018-09-18T13:20:57Z")

</div>

This is really helpful for novice to Julia like me.

---

<div class="post-metadata">

**Author:** ![Mikkel-Holm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikkel-holm/32/5502_2.png) [@Mikkel-Holm](https://discourse.julialang.org/u/Mikkel-Holm)\
**Post date:** [September 25, 2018, 10:04am UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/8 "2018-09-25T10:04:30Z")

</div>

Plug and play snippet for train / test data sample split if your data is in the format of a **multi-dimensional array**.  
_@Evizero has a snippet above if your data is in DataFrame format._

```julia
"""
partitionTrainTest(data, y_column, train_perc = 0.7).
\n Returns X_train, y_train, X_test, y_test given inputs data (multi-dimensional array), 
\n train_perc (percent of data in train sample) and y_column (column location of y values).
"""
function partitionTrainTest(data, y_column, train_perc = 0.7)
    n = size(data,1)
    idx = shuffle(1:n)
    train_idx = view(idx, 1:floor(Int, at*n))
    test_idx = view(idx, (floor(Int, at*n)+1):n)
    data[train_idx, 1:end .!=y_column], data[train_idx, y_column], data[test_idx, 1:end .!=y_column], data[test_idx, y_column]
end

```

---

<div class="post-metadata">

**Author:** ![Geoffrey\_Anderson](https://avatars.discourse-cdn.com/v4/letter/g/94ad74/32.png) [@Geoffrey\_Anderson](https://discourse.julialang.org/u/Geoffrey_Anderson)\
**Post date:** [October 1, 2018, 6:59pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/9 "2018-10-01T18:59:32Z")

</div>

A 3-partition split is more useful than 2-partition if you ever want to evaluate your model which is convincing to anyone else. It’s not convincing on a 2-partition set that your predictions are any good.  
Furthermore sometimes even 4 or more partitions are needed depending (typically seen in deep learning models). This is because new data is needed for unbiased error estimation of the final model after bias and variance errors have been seemingly learned to be minimized by the learning algorithm and your hyperparameters such as L2 regularization. You cannot use the “test” partition of a 2-partition split on which you chose the best L2 or polynomial degree or whatever, to evaluate your model honestly (without bias).

How would you make a 3 partition split: train/dev/test?

Question 2: Is the implementation going to work on big data? We don’t generally have the RAM to load an entire dataset once, never mind twice, so please use care to avoid making a duplicate memory allocation of original data (monolithic) as well as partitioned data (2, 3, or more partitions as I described).

Thank you!

---

<div class="post-metadata">

**Author:** ![Geoffrey\_Anderson](https://avatars.discourse-cdn.com/v4/letter/g/94ad74/32.png) [@Geoffrey\_Anderson](https://discourse.julialang.org/u/Geoffrey_Anderson)\
**Post date:** [October 1, 2018, 7:02pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/10 "2018-10-01T19:02:04Z")

</div>

> [@Evizero](#):
>
> partitionTrainTest

I am frankly worried due to your choice of function name baked in here, that only 2 partitions are going to be possible when that documentation is complete, i.e., train, test. Are you guys going to allow 3 partitions or more (preferably more)? What does the library author plan to name the function that supports slicing into more than 2 data partitions? (See my other post for more background on partitioning.)

---

<div class="post-metadata">

**Author:** ![platawiec](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/platawiec/32/31914_2.png) [@platawiec](https://discourse.julialang.org/u/platawiec)\
**Post date:** [October 1, 2018, 9:14pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/11 "2018-10-01T21:14:17Z")

</div>

Check out MLDataUtils.jl, particularly `splitobs`. See [this page](https://mldatautilsjl.readthedocs.io/en/latest/data/pattern.html#splitting-into-train-test) for a 3-partition example.

The post you are replying to is almost 2 years old, and there’s been some progress since then 😉

---

<div class="post-metadata">

**Author:** ![Evizero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/evizero/32/10118_2.png) [@Evizero](https://discourse.julialang.org/u/Evizero)\
**Post date:** [October 1, 2018, 9:51pm UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/12 "2018-10-01T21:51:11Z")

</div>

> [@platawiec](#):
>
> Check out MLDataUtils.jl, particularly `splitobs` . See [this page](https://mldatautilsjl.readthedocs.io/en/latest/data/pattern.html#splitting-into-train-test) for a 3-partition example.
> 
> The post you are replying to is almost 2 years old, and there’s been some progress since then

Thank you for the reply and for reading the docs

I would mark your post as an answer to this ancient thread if I could

---

<div class="post-metadata">

**Author:** ![microgold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/microgold/32/5404_2.png) [@microgold](https://discourse.julialang.org/u/microgold)\
**Post date:** [March 20, 2020, 1:09am UTC](https://discourse.julialang.org/t/simple-tool-for-train-test-split/473/13 "2020-03-20T01:09:54Z")

</div>

The third set was in a separate file, but i do see your point!
