# DataFrame to multidimensional array

**URL:** <https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [December 15, 2022, 8:54am UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686 "2022-12-15T08:54:12Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![m\_pink](https://avatars.discourse-cdn.com/v4/letter/m/c37758/32.png) [@m\_pink](https://discourse.julialang.org/u/m_pink)\
**Post date:** [December 15, 2022, 8:54am UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/1 "2022-12-15T08:54:12Z")

</div>

Hey,

Suppose I have a Dataframe that corresponds to a multidimensional function. Something like this:  
`df = DataFrame([(a = x, b = y, c = z,d = x+y+z) for x in 1:6 for y in 1:4 for z in 1:2])`

What is the best way to transform it into the following multidimensional array  
`[x+y+z for x in 1:6, y in 1:4 , z in 1:2]`

Currently I’m doing it with `groupby` but it gets complicated with higher dimensions (I’m a Matlab user, so I used to work with multidimensional arrays).

Thank you

---

<div class="post-metadata">

**Author:** ![rmsmsgood](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmsmsgood/32/20544_2.png) [@rmsmsgood](https://discourse.julialang.org/u/rmsmsgood)\
**Post date:** [December 15, 2022, 10:46am UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/2 "2022-12-15T10:46:20Z")

</div>

Are x,y,z integer?

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [December 15, 2022, 11:40am UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/3 "2022-12-15T11:40:43Z")

</div>

I assume you know the sizes of x, y and z? Inferring those from the vectors first would be a more annoying step.

```julia
df = DataFrame([(a = x, b = y, c = z,d = x+y+z) for x in 1:6 for y in 1:4 for z in 1:2])

arr = permutedims(reshape(copy(df.d), (2, 4, 6)), (3, 2, 1))

```

This gives:

```julia
6×4×2 Array{Int64, 3}:
[:, :, 1] =
 3 4 5 6
 4 5 6 7
 5 6 7 8
 6 7 8 9
 7 8 9 10
 8 9 10 11

[:, :, 2] =
 4 5 6 7
 5 6 7 8
 6 7 8 9
 7 8 9 10
 8 9 10 11
 9 10 11 12

julia> arr == [x+y+z for x in 1:6, y in 1:4 , z in 1:2]
true

```

Note that `for x in 1:6 for y in 1:4 for z in 1:2` has exactly the opposite order of dimensions than `for x in 1:6, y in 1:4, z in 1:2` which is why the `permutedims` is needed.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 15, 2022, 1:04pm UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/4 "2022-12-15T13:04:07Z")

</div>

Here is an alternative using TensorCast.jl:

```julia
using TensorCast
@cast v[i,j,k] := copy(df.d)[k⊗j⊗i] (i ∈ 1:6, j ∈ 1:4, k ∈ 1:2)

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [December 15, 2022, 9:40pm UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/5 "2022-12-15T21:40:08Z")

</div>

Depending on how you obtain the data in the first place, you may refactor that process to return a multidimensional array instead of a dataframe. Arrays are indeed easy and convenient to use in julia, and they are more general.

But for this particular operation, there’s a nice table → multi dim array conversion function in `AxisKeys.jl`:

```julia
julia> using AxisKeys

julia> wrapdims(df, :d, :a, :b, :c)
3-dimensional KeyedArray(NamedDimsArray(...)) with keys:
↓ a ∈ 6-element Vector{Int64}
→ b ∈ 4-element Vector{Int64}
◪ c ∈ 2-element Vector{Int64}
And data, 6×4×2 Array{Int64, 3}:
[:, :, 1] ~ (:, :, 1):
      (1) (2) (3) (4)
 (1) 3 4 5 6
 (2) 4 5 6 7
 (3) 5 6 7 8
 (4) 6 7 8 9
 (5) 7 8 9 10
 (6) 8 9 10 11

[:, :, 2] ~ (:, :, 2):
      (1) (2) (3) (4)
 (1) 4 5 6 7
 (2) 5 6 7 8
 (3) 6 7 8 9
 (4) 7 8 9 10
 (5) 8 9 10 11
 (6) 9 10 11 12

```

It would even work with non-consecutive or non-numeric `x`, `y`, `z` values.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 15, 2022, 9:58pm UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/6 "2022-12-15T21:58:52Z")

</div>

The performance of `wrapdims()` on this specific example seems to be way subpar.

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [February 26, 2024, 7:10am UTC](https://discourse.julialang.org/t/dataframe-to-multidimensional-array/91686/7 "2024-02-26T07:10:09Z")

</div>

Can this be combined with DataFrames groupby or @by to wrap a sub-set of the dimensions, to produce a DataFrame where one of the columns is a KeyedArray.

The DataFrames guys are working on `nest` and `unnest` which would make this easier but they aren’t released yet.

> <https://github.com/JuliaData/DataFrames.jl/pull/3258>
>
> Fixes
> https://github.com/JuliaData/DataFrames.jl/issues/3005
> https://github.co…m/JuliaData/DataFrames.jl/issues/2890
> https://github.com/JuliaData/DataFrames.jl/issues/3116
> https://github.com/JuliaData/DataFrames.jl/issues/2767
> 
> The PR adds \`nest\` and \`unnest\` and introduces \`scalar\` kwarg to \`flatten\` (which is needed in \`unnest\`.
> 
> \`flatten\` is ready for review.
> 
> For \`nest\` and \`unnest\` requires discussion if we like the proposed API (they work, but maybe we will decide to change API).
> 
> Some important decisions I propose:
> \* \`nest\` only works on \`GroupedDataFrame\` (the reason is to avoid complexity of group order specification); nesting is done always to \`DataFrame\` (to keep things simple); another not easy decision is syntax I proposed \`\[:x, :y\] =\> :z\` which means that columns \`:x\` and \`:y\` should be nested and stored in column \`:z\` (but I would like to confirm that we find it intuitive, as syntax \`:z =\> \[:x, :y\]\` also could be advocated for).
> \* \`unnest\` supports both tables (e.g. \`DataFrame\`) and rows (e.g. \`Tables.AbstractRow\`) and has two options: \`flatten=true\`, when rows of the nested columns are flattened, and \`flatten=false\` (when they are left as is - this is probably useful, if we work with rows)
> 
> TODO:
> 
> \- add metadata
> \- write tests
> \- update manual
> 
> CC @nalimilan @pdeffebach @jariji
