# How to use \`ordered\` keyword equals to \`true\` with \`StatsBase.sample()\` function when sampling 2 or more dimensional array from a population?

**URL:** <https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796>\
**Category:** Modelling & Simulations\
**Created:** [March 23, 2021, 3:59pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796 "2021-03-23T15:59:11Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![shubh\_b](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shubh_b/32/20697_2.png) [@shubh\_b](https://discourse.julialang.org/u/shubh_b)\
**Post date:** [March 23, 2021, 3:59pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/1 "2021-03-23T15:59:11Z")

</div>

I was using `sample()` function from `StatsBase` module for sampling with replacement from a population `a` and sampling probabilities weight vector `wv` as below:

```julia
a = [31, 12, 29, 9, 15]
wv = weights([0.1, 0.4, 0.2, 0.1, 0.2])
sample(a, wv, (3, 4), ordered=false)

```

It provides a random matrix that looks like:

```julia
3×4 Array{Int64,2}:
 31 12 29 29    
 15 9 9 31    
 31 9 15 15

```

But when I write the code as:

`sample(a, wv, (3, 4), ordered=true)`

instead of providing any ordered (by row or column) random matrix, it throws the error:

`ERROR: UndefKeywordError: keyword argument dims not assigned`

How to use the keyword argument `ordered=true` here for sampling a 2 or more dimensional array from a population array?

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [March 23, 2021, 4:13pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/2 "2021-03-23T16:13:58Z")

</div>

That error message could definitely be made better, but I think the reason why your example doesn’t work is because there isn’t a clear way to order the values of a matrix. For example, if your sampled values are `(1, 2, 3, 4)`, is the correct order for a 2x2 matrix this,

```julia
1 2
3 4

```

or this,

```julia
1 3
2 4

```

or even this,

```julia
4 3
2 1

```

What you could do is sample a vector and then reshape it into a matrix:

```julia
a = [31, 12, 29, 9, 15]
wv = weights([0.1, 0.4, 0.2, 0.1, 0.2])
x = sample(a, wv, 12, ordered=true)

```

```julia
julia> reshape(x, 3, 4)
3×4 Array{Int64,2}:
 12 12 15 29
 12 15 15 29
 12 15 15 29

```

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [March 23, 2021, 4:27pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/3 "2021-03-23T16:27:23Z")

</div>

I opened a Github issue to improve the error message:

[https://github.com/JuliaStats/StatsBase.jl/issues/676](https://github.com/JuliaStats/StatsBase.jl/issues/676)

---

<div class="post-metadata">

**Author:** ![shubh\_b](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shubh_b/32/20697_2.png) [@shubh\_b](https://discourse.julialang.org/u/shubh_b)\
**Post date:** [March 23, 2021, 4:31pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/4 "2021-03-23T16:31:31Z")

</div>

Thanks for your so quick reply and sharing your thoughts.

Similar thing I also tried. This is also not a very straightforward way. If we write:

```julia
sort(sample(a, wv, (3, 4)), dims=2) # For ordering column-wise.
3×4 Array{Int64,2}:
 12 15 15 29
  9 9 12 31
 12 12 15 15

sort(sample(a, wv, (3, 4)), dims=1) # For ordering row-wise.
3×4 Array{Int64,2}:
 12 9 12 12
 12 15 12 15
 29 29 29 31

```

Obviously, in these ways we are not getting the desired output directly from the `sample()` function.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [March 23, 2021, 4:35pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/5 "2021-03-23T16:35:44Z")

</div>

A basic issue here is that there is a bug in `sample(a, ordered=true, replace=true)`: [incorrect ordering for `sample(replace=true, ordered=true)` · Issue #675 · JuliaStats/StatsBase.jl · GitHub](https://github.com/JuliaStats/StatsBase.jl/issues/675)

You’ll see that if you sample a 1d output it gives a non-sequential order (i.e. not matching the ordering in `a`):

```julia
julia> sample(a, wv, 3, ordered=true)
3-element Array{Int64,1}:
  9
 12
 29

```

> [@CameronBieganek](#):
>
> That error message could definitely be made better, but I think the reason why your example doesn’t work is because there isn’t a clear way to order the values of a matrix.

The natural way in Julia would be iteration order, i.e. column-major order.

Except for the abovementioned bug you could do this via `reshape`:

```julia
julia> reshape(sample(a, wv, 12, ordered=true), 3,4)
3×4 Array{Int64,2}:
  9 12 15 15
 12 12 15 15
 12 12 15 29

```

A workaround to get a correct ordering would be to sample indices:

```julia
julia> reshape(a[sample(1:length(a), wv, 12, ordered=true)], 3,4)
3×4 Array{Int64,2}:
 31 12 29 15
 31 29 9 15
 12 29 15 15

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [March 23, 2021, 4:45pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/6 "2021-03-23T16:45:37Z")

</div>

> [@shubh\_b](#):
>
> `sort(sample(a, wv, (3, 4)), dims=2) # For ordering column-wise.`

(This is not the same thing as ordered/sequential sampling. Which is fine, if that’s what you want.)

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [March 23, 2021, 4:56pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/7 "2021-03-23T16:56:21Z")

</div>

> [@stevengj](#):
>
> You’ll see that if you sample a 1d output it gives a non-sequential order (i.e. not matching the ordering in `a` )

Ah, I see what you mean. I just assumed that the `ordered` keyword argument meant that the output samples should be sorted (e.g., lexicographically). However, from a quick google search I can’t find a reference for the type of sorting you refer to, where the output order respects the order in the original vector. The closest I’ve found is this:

> **[Sampling in order](https://en.wikipedia.org/wiki/Sampling_in_order)**
>
> In statistics, some Monte Carlo methods require independent observations in a sample to be drawn from a one-dimensional distribution in sorted order. In other words, all n order statistics are needed from the n observations in a sample. The naive method performs a sort and takes O(n log n) time. There are also O(n) algorithms which are better suited for large n. The special case of drawing n sorted observations from the uniform distribution on \[0,1\] is equivalent to drawing from the uniform distr...

which describes the case where you sample and then sort the samples (for the naive algorithm).

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [March 23, 2021, 5:58pm UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/8 "2021-03-23T17:58:33Z")

</div>

> [@CameronBieganek](#):
>
> Ah, I see what you mean. I just assumed that the `ordered` keyword argument meant that the output samples should be sorted (e.g., lexicographically). However, from a quick google search I can’t find a reference for the type of sorting you refer to, where the output order respects the order in the original vector.

For example, [Vitter (1984)](https://dl.acm.org/doi/abs/10.1145/358105.893) defines “sequential” sampling this way, and the StatsBase docs define “ordered” as equivalent to “sequential”.

Moreover, `StatsBase.sample` uses this definition for the `replace=false` case, so it makes no sense to use a completely different definition of “ordered” for sampling with replacement.

Finally, sequential sampling as Vitter defined it is strictly more general. If you want sorted results from `sample(a, x, ordered=true)`, you can simply `sort!(a)` first in any ordering you want. (And the definition you are using fails completely if the elements of `a` do not have an `isless` ordering.)

---

<div class="post-metadata">

**Author:** ![shubh\_b](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shubh_b/32/20697_2.png) [@shubh\_b](https://discourse.julialang.org/u/shubh_b)\
**Post date:** [March 24, 2021, 4:05am UTC](https://discourse.julialang.org/t/how-to-use-ordered-keyword-equals-to-true-with-statsbase-sample-function-when-sampling-2-or-more-dimensional-array-from-a-population/57796/9 "2021-03-24T04:05:01Z")

</div>

Yes, that’s true. What I tried, is not the same thing as we see in sequential sampling.
