# Transform! to destructure NamedTuple into columns

**URL:** <https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991>\
**Category:** General Usage\
**Tags:** question, dataframes\
**Created:** [January 21, 2022, 3:26pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991 "2022-01-21T15:26:36Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [January 21, 2022, 3:26pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/1 "2022-01-21T15:26:36Z")

</div>

Hi all,

I would like to destructure a DataFrame with NamedTuples into separate columns, using the keys as column names. For a similar problem with an array, I used something like `transform!(df, :col => :identity => new_names)`. Is there a way I can take this:

```julia
2×2 DataFrame
 Row │ x y              
     │ Float64 NamedTup…      
─────┼──────────────────────────
   1 │ 0.222043 (a = 1, b = 2)
   2 │ 0.72646 (a = 3, b = 5)

```

and obtain this:

```julia
2×4 DataFrame
 Row │ x y a b     
     │ Float64 NamedTup… Int64 Int64 
─────┼────────────────────────────────────────
   1 │ 0.222043 (a = 1, b = 2) 1 2
   2 │ 0.72646 (a = 3, b = 5) 3 5

```

Thanks!

**MWE**

```julia
using DataFrames

df = DataFrame(x = rand(2), y =[(a=1,b=2),(a=3,b=5)])

df_new = DataFrame(x = df.x, y = df.y, a = [1,3], b = [2,5])

```

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [January 21, 2022, 3:32pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/2 "2022-01-21T15:32:43Z")

</div>

You can use `AsTable`:

```julia
julia> transform(df, :y => AsTable)
2×4 DataFrame
 Row │ x y a b     
     │ Float64 NamedTup… Int64 Int64 
─────┼────────────────────────────────────────
   1 │ 0.459213 (a = 1, b = 2) 1 2
   2 │ 0.241038 (a = 3, b = 5) 3 5

```

---

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [January 21, 2022, 3:45pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/3 "2022-01-21T15:45:12Z")

</div>

Very nice. Thanks!

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [January 21, 2022, 4:32pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/4 "2022-01-21T16:32:46Z")

</div>

> [@sijo](#):
>
> `transform(df, :y => AsTable)`

This seems to be equivalent to:

```julia
hcat(df, DataFrame(df.y))

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [January 21, 2022, 4:36pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/5 "2022-01-21T16:36:13Z")

</div>

Yes, but it will do more allocations (which is a minor issue but still might be relevant occasionally).

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [January 21, 2022, 4:41pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/6 "2022-01-21T16:41:30Z")

</div>

> [@bkamins](#):
>
> will do more allocations

Obviously I am doing something wrong, but for the small OP example I actually see less allocations?

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [January 21, 2022, 5:03pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/7 "2022-01-21T17:03:29Z")

</div>

Ah yes I also see less allocations with your solution… @bkamins ?

By the way, since `[a b]` is fancy syntax for `hcat(a, b)` you can also write

```julia
[df DataFrame(df.y)]

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [January 21, 2022, 5:15pm UTC](https://discourse.julialang.org/t/transform-to-destructure-namedtuple-into-columns/74991/8 "2022-01-21T17:15:42Z")

</div>

Ah - you are right:

```julia
julia> df = repeat(DataFrame(x=1, y=(a=1,b=2)), 10^8);

julia> @time transform(df, :y => AsTable);
  5.982615 seconds (200.00 M allocations: 9.686 GiB, 7.61% gc time)

julia> @time [df DataFrame(df.y)];
  1.345369 seconds (61 allocations: 5.215 GiB, 6.20% gc time)

```

This means that I need to optimize the internals of `transform` 😄.

This guarantees that there is no aliasing between source and target and at the same time that we do not do unnecessary allocations:

```julia
julia> @time hcat(copy(df), DataFrame(df.y), copycols=false);
  1.231519 seconds (67 allocations: 3.725 GiB, 35.97% gc time)

```
