# Sample rows of a 3 dimensional array without replacement

**URL:** <https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116>\
**Category:** General Usage\
**Tags:** question, statistics, arrays, cartesianindices\
**Created:** [November 7, 2021, 11:30pm UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116 "2021-11-07T23:30:50Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [November 7, 2021, 11:30pm UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/1 "2021-11-07T23:30:51Z")

</div>

I would like to sample rows from a three dimensional array without replacement. For example, if I have

`x = rand(10, 5, 3)`

a random set of indices might look like `{(2,2), (9,3)}`, where the first position is the index of the first dimension and the second position is the index of the third dimension. If I wanted a point, I could use linear indexing. However, it is a bit trickier because I need to ignore the second row. Is there a good way to do this in Julia?

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [November 8, 2021, 12:12am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/2 "2021-11-08T00:12:25Z")

</div>

Maybe something like:

```julia
using StatsBase
x = rand(10, 5, 3)
@views sample(CartesianIndices(x[:,1,:]), 3; replace=false)

3-element Vector{CartesianIndex{2}}:
 CartesianIndex(1, 3)
 CartesianIndex(5, 2)
 CartesianIndex(4, 2)

```

---

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [November 8, 2021, 12:19am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/3 "2021-11-08T00:19:29Z")

</div>

Awesome. Thanks!

---

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [November 8, 2021, 12:35am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/4 "2021-11-08T00:35:12Z")

</div>

Do you happen to know of a more efficient solution for large arrays?

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [November 8, 2021, 5:46am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/6 "2021-11-08T05:46:16Z")

</div>

A simple one-liner that copies:

```julia
using StatsBase, SplitApplyCombine

sample(splitdims(x, (1, 3)), 20; replace=false)

```

More efficient, doesn’t copy array content:

```julia
splitted = splitdimsview(x, (1, 3))
I = sample(CartesianIndices(splitted), 20; replace=false)
@view splitted[I] # 20-element view(...)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [November 8, 2021, 7:27am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/7 "2021-11-08T07:27:19Z")

</div>

Fyi, see herein some benchmarks:

```julia
using StatsBase
f(x,n) = @views sample(CartesianIndices(x[:,1,:]), n; replace=false)

using SplitApplyCombine
g(x,n) = sample(splitdims(x, (1, 3)), n; replace=false)

function h(x,n)
    splitted = splitdimsview(x, (1, 3))
    I = sample(CartesianIndices(splitted), n; replace=false)
    @view splitted[I]
end

n = 1000
x = rand(100, 50, 30)

using BenchmarkTools
@btime f($x,$n) # 8.9 μs (3 allocations: 39.23 KiB)
@btime g($x,$n) # 205.7 μs (3005 allocations: 1.47 MiB)
@btime h($x,$n) # 9.4 μs (3 allocations: 39.23 KiB)

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [November 8, 2021, 8:35am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/8 "2021-11-08T08:35:16Z")

</div>

This benchmark comparison looks totally reasonable, with the expected order between functions.  
Note that `f` returns indices, while `g` and `h` return array elements. Also, I don’t think these 2d indices from `f()` can be used to directly index into 3d `x`…

---

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [November 8, 2021, 8:54am UTC](https://discourse.julialang.org/t/sample-rows-of-a-3-dimensional-array-without-replacement/71116/9 "2021-11-08T08:54:09Z")

</div>

Thank you both for the ideas!
