# Dataframe to nested tuple?

**URL:** <https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182>\
**Category:** General Usage\
**Tags:** question, dataframes, namedtuple\
**Created:** [December 3, 2020, 11:28am UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182 "2020-12-03T11:28:21Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [December 3, 2020, 11:28am UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/1 "2020-12-03T11:28:21Z")

</div>

Id like to transform this:

```julia
DataFrame( GroupLetter = ['A','A','B','B'] , GroupID = [1,1,2,2], Col1 = [1,2,3,4], Col2 = [5,6,7,8] )

 | GroupLetter │ GroupID │ Col1 │ Col2 │
 │ 'A' │ 1 │ 1 │ 5 │
 │ 'A' │ 1 │ 2 │ 6 │
 │ 'B' │ 2 │ 3 │ 7 │
 │ 'B' │ 2 │ 4 │ 8 │

```

Into this:

```julia
data = ( 
    A = (GroupID = 1, DF = DataFrame(Col1 = [1,2], Col2 = [5,6]) ),
    B = (GroupID = 2, DF = DataFrame(Col1 = [3,4], Col2 = [7,8]) )
)

```

So I can write

```julia
data.A.GroupID 

```

and

```julia
data.A.DF

```

giving

```julia
│ Col1 │ Col2 │
│ 1 │ 5 │
│ 2 │ 6 │

```

Is there an easy way to do this?  
And, can the nested tuple structure be written to and from a file ?

---

<div class="post-metadata">

**Author:** ![fbanning](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fbanning/32/14972_2.png) [@fbanning](https://discourse.julialang.org/u/fbanning)\
**Post date:** [December 3, 2020, 11:48am UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/2 "2020-12-03T11:48:38Z")

</div>

```julia
julia> data = DataFrame( GroupLetter = ['A','A','B','B'] , GroupID = [1,1,2,2], Col1 = [1,2,3,4], Col2 = [5,6,7,8] )
4×4 DataFrame
│ Row │ GroupLetter │ GroupID │ Col1 │ Col2 │
│ │ Char │ Int64 │ Int64 │ Int64 │
├─────┼─────────────┼─────────┼───────┼───────┤
│ 1 │ 'A' │ 1 │ 1 │ 5 │
│ 2 │ 'A' │ 1 │ 2 │ 6 │
│ 3 │ 'B' │ 2 │ 3 │ 7 │
│ 4 │ 'B' │ 2 │ 4 │ 8 │

julia> groupby(data, ["GroupLetter", "GroupID"])
GroupedDataFrame with 2 groups based on keys: GroupLetter, GroupID
First Group (2 rows): GroupLetter = 'A', GroupID = 1
│ Row │ GroupLetter │ GroupID │ Col1 │ Col2 │
│ │ Char │ Int64 │ Int64 │ Int64 │
├─────┼─────────────┼─────────┼───────┼───────┤
│ 1 │ 'A' │ 1 │ 1 │ 5 │
│ 2 │ 'A' │ 1 │ 2 │ 6 │
⋮
Last Group (2 rows): GroupLetter = 'B', GroupID = 2
│ Row │ GroupLetter │ GroupID │ Col1 │ Col2 │
│ │ Char │ Int64 │ Int64 │ Int64 │
├─────┼─────────────┼─────────┼───────┼───────┤
│ 1 │ 'B' │ 2 │ 3 │ 7 │
│ 2 │ 'B' │ 2 │ 4 │ 8 │

julia> gdf[(GroupLetter = 'A', GroupID = 1)]
2×4 SubDataFrame
│ Row │ GroupLetter │ GroupID │ Col1 │ Col2 │
│ │ Char │ Int64 │ Int64 │ Int64 │
├─────┼─────────────┼─────────┼───────┼───────┤
│ 1 │ 'A' │ 1 │ 1 │ 5 │
│ 2 │ 'A' │ 1 │ 2 │ 6 │

```

Not quite what you asked for but maybe close enough that you like it.

Edit: Remember when doing such a thing that piping (conveniently via Pipe.jl or Chain.jl) is always more ~~performant~~ concise.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [December 3, 2020, 1:52pm UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/3 "2020-12-03T13:52:22Z")

</div>

> [@fbanning](#):
>
> Edit: Remember when doing such a thing that piping (conveniently via Pipe.jl or Chain.jl) is always more performant.

I don’t mean to derail this but piping should not be any more or less performant than other ways of writing the same code; it’s just syntax.

---

<div class="post-metadata">

**Author:** ![fbanning](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fbanning/32/14972_2.png) [@fbanning](https://discourse.julialang.org/u/fbanning)\
**Post date:** [December 3, 2020, 2:39pm UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/4 "2020-12-03T14:39:10Z")

</div>

Bad wording, sorry.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 3, 2020, 2:44pm UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/5 "2020-12-03T14:44:38Z")

</div>

> [@fbanning](#):
>
> `gdf[(GroupLetter = 'A', GroupID = 1)]`

or `gdf[('A', 1)]` if you want to avoid passing the grouping column names.

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [December 3, 2020, 11:14pm UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/6 "2020-12-03T23:14:10Z")

</div>

Thank you both.  
It seems if there is only one grouping column you need a comma in the tuple

gdf = groupby(data, :GroupLetter)

gdf[( :A, )]

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [December 3, 2020, 11:29pm UTC](https://discourse.julialang.org/t/dataframe-to-nested-tuple/51182/7 "2020-12-03T23:29:13Z")

</div>

Yeah, `tuple(x)` looks slightly less clumsy imo
