# Stratified and weighted sampling in dataframes

**URL:** <https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565>\
**Category:** Statistics\
**Tags:** question, dataframes\
**Created:** [March 8, 2022, 7:13am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565 "2022-03-08T07:13:08Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![George\_Githinji](https://avatars.discourse-cdn.com/v4/letter/g/e68b1a/32.png) [@George\_Githinji](https://discourse.julialang.org/u/George_Githinji)\
**Post date:** [March 8, 2022, 7:13am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/1 "2022-03-08T07:13:08Z")

</div>

I have a function that can take a random sample from a dataframe. However I ran into an issue when working with grouped a dataframe in which case I have grouped the data by several columns and I would like to sample within these groups. I cannot find an obvious way to go about or a detailed tutorial.  
So I have the following,

```julia
function take_a_sample(df, size)
    df[sample(axes(df, 1), size; replace = false, ordered = true), :]
end

```

Where `df` is the dataframe and `size` is the number of samples i would like. This works fine with plain dataframe but then i need to group and sample within groups  
`df_grouped = groupby(df,[:country,:date,:location])`

`take_a_sample(df_grouped,100)`

I get `ERROR: ArgumentError: GroupedDataFrame requires a single index`

How can I get around sampling uniformly across multiple groups in a grouped dataframe.  
Thanks.

---

<div class="post-metadata">

**Author:** ![DorianT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/doriant/32/24897_2.png) [@DorianT](https://discourse.julialang.org/u/DorianT)\
**Post date:** [March 8, 2022, 7:55am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/2 "2022-03-08T07:55:24Z")

</div>

Without trying a MWE I think the problem is that `df` in the case where you get an error is a grouped data frame, but your function `take_a_sample` constructs an array of indicies which you then try to index the grouped data frame with (but I believe when indexing a grouped data frame the first index should be a scalar indexing the sub data frame you want to select). One way you can fix this should be to call `take_a_sample` from a loop that goes over the sub data frames.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 8, 2022, 7:56am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/3 "2022-03-08T07:56:57Z")

</div>

```julia
take_a_sample(df::AbstractDataFrame, size) =
    df[sample(axes(df, 1), size; replace = false, ordered = true), :]
take_a_sample(gdf::GroupedDataFrame, size) =
    combine(gdf, x -> take_a_sample(x, size))

```

is simplest

---

<div class="post-metadata">

**Author:** ![George\_Githinji](https://avatars.discourse-cdn.com/v4/letter/g/e68b1a/32.png) [@George\_Githinji](https://discourse.julialang.org/u/George_Githinji)\
**Post date:** [March 8, 2022, 9:39am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/4 "2022-03-08T09:39:32Z")

</div>

Many thanks - it looks like there is small typo in your response ( `df_grouped` should be `gdf`).

I ran into an issue when I tried to sample with a size \> 1 e.g. t`ake_a_sample (gdf, 10)` with the grouped dataframe.

`

`ERROR: LoadError: Cannot draw more samples without replacement.`

```julia
 nested task error: Cannot draw more samples without replacement.
    Stacktrace:

```

`

Also if sample size is 1 it returns a dataframe of the same number of rows as the grouping hierarchy. I think this would be expected because it sampling once per group? Sorry if i was not clear, but would want to sample multiply times within each group (from the grouped dataframe)

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 8, 2022, 9:56am UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/5 "2022-03-08T09:56:57Z")

</div>

> [@George\_Githinji](#):
>
> I ran into an issue when I tried to sample with a size \> 1 e.g. t `ake_a_sample (gdf, 10)` with the grouped dataframe.

This means that some groups in your data frame have less than 10 rows. This is expected if you want to do sampling without replacement

> [@George\_Githinji](#):
>
> Also if sample size is 1 it returns a dataframe of the same number of rows as the grouping hierarchy. I think this would be expected because it sampling once per group?

This is also expected and correct.

> [@George\_Githinji](#):
>
> Sorry if i was not clear, but would want to sample multiply times within each group (from the grouped dataframe)

You mean you want to do sampling with replacement? Then use `replace=true`.

---

<div class="post-metadata">

**Author:** ![George\_Githinji](https://avatars.discourse-cdn.com/v4/letter/g/e68b1a/32.png) [@George\_Githinji](https://discourse.julialang.org/u/George_Githinji)\
**Post date:** [March 8, 2022, 5:27pm UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/6 "2022-03-08T17:27:44Z")

</div>

Thank you so much for the detailed explanation and in addition to the quick response! I ideally want to sample within each group x number of times but if the group size is less than the samples then you would want to pick all the samples in the group and continue without an error. Sampling with replacement means that I draw the same rows/entries multiple times i.e. I would get a duplicated rows. This is OK when the number of samples is less than the groups ( I could then look for unique rows by removing duplicates). I would however not wish to sample the same entry if there are additional non-unique rows or entries, thus i would want to keep `replace=false` .

---

<div class="post-metadata">

**Author:** ![George\_Githinji](https://avatars.discourse-cdn.com/v4/letter/g/e68b1a/32.png) [@George\_Githinji](https://discourse.julialang.org/u/George_Githinji)\
**Post date:** [March 8, 2022, 5:36pm UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/7 "2022-03-08T17:36:15Z")

</div>

@bkamins Slightly off-topic but do you have some recommended tutorials or books or a list of howtos that you could recommend ?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 8, 2022, 7:11pm UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/8 "2022-03-08T19:11:06Z")

</div>

> [@George\_Githinji](#):
>
> I ideally want to sample within each group x number of times but if the group size is less than the samples then you would want to pick all the samples in the group and continue without an error.

```julia
take_a_sample(df::AbstractDataFrame, size) =
    df[sample(axes(df, 1), min(size, nrow(df)); replace = false, ordered = true), :]

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 8, 2022, 7:11pm UTC](https://discourse.julialang.org/t/stratified-and-weighted-sampling-in-dataframes/77565/9 "2022-03-08T19:11:37Z")

</div>

> [@George\_Githinji](#):
>
> Slightly off-topic but do you have some recommended tutorials or books or a list of howtos that you could recommend ?

On what topic? If you mean DataFrames.jl then please check out [Introduction · DataFrames.jl](https://dataframes.juliadata.org/stable/).
