# Construct DataFrame From Uneven Named Tuples

**URL:** <https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [August 19, 2023, 1:59am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970 "2023-08-19T01:59:55Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![quietlight](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quietlight/32/35187_2.png) [@quietlight](https://discourse.julialang.org/u/quietlight)\
**Post date:** [August 19, 2023, 1:59am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/1 "2023-08-19T01:59:55Z")

</div>

I am trying to construct a DataFrame from an array of named tuples where some key:value pairs may be missing from some tuples.

They come from querying Airtable using Airtable,jl. If you have missing values in your table, they are not present in the data returned by the airtable API.

This is what an array may look like:

data = [(Number = 4, Name = “Abc Efgg”, Address = “48 Mont Rd”),  
(Number = 6, Name = “Ruf Sly”, Address = “19A Keke ava”),  
(Number = 10, Name = “Jack Bog”),  
(Number = 5, Name = “Gid Hoo”, Address = “120 Mut Street”)]

I seem to be struggling to find a simple solution. Any ideas welcome.

Regards  
David

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [August 19, 2023, 2:42am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/2 "2023-08-19T02:42:45Z")

</div>

One way would be:

```julia
cols = union(keys.(data)...)
df = DataFrame([c => get.(data,c, missing) for c in cols]...)

```

which gives:

```julia
4×3 DataFrame
 Row │ Number Name Address        
     │ Int64 String String?        
─────┼──────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 10 Jack Bog missing        
   4 │ 5 Gid Hoo 120 Mut Street

```

But the real question is how the annoying Unicode double quotes ““” managed to get into the OP 😬

---

<div class="post-metadata">

**Author:** ![quietlight](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quietlight/32/35187_2.png) [@quietlight](https://discourse.julialang.org/u/quietlight)\
**Post date:** [August 19, 2023, 3:14am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/3 "2023-08-19T03:14:50Z")

</div>

> [@Dan](#):
>
> `cols = union(keys.(data)...)`

Thank you for your kind reply. get! I forgot about that.

Regards  
David

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 19, 2023, 6:52am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/4 "2023-08-19T06:52:06Z")

</div>

Probably this scheme is valid also without the intervention of the foraeach function (that is, using only the named tuples and some splatting/broadcasting), but I haven’t found the way yet

```julia
df=DataFrame()
foreach(d->push!(df,d,cols=:union),data)

```

i would have expected that push! worked the same way in the following two cases

```julia
push!([1,2],3,4)

push!(df,data...,cols=:union)

```

Oddly enough this way, it works

```julia
push!.([df],data,cols=:union)

push!.([df],data,cols=:union)[1]

```

trying append!(…,cols=:union) I would have expected it to handle the missing field

```julia
julia> append!(df,data,cols=:union)
ERROR: type NamedTuple has no field Address

```

which dictrowtable does instead

```julia
append!(df,Tables.dictrowtable(data))

# or better

julia> DataFrame(Tables.dictrowtable(data))
4×3 DataFrame
 Row │ Number Name Address        
     │ Int64 String String?
─────┼──────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 10 Jack Bog missing
   4 │ 5 Gid Hoo 120 Mut Street

```

---

<div class="post-metadata">

**Author:** ![quietlight](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quietlight/32/35187_2.png) [@quietlight](https://discourse.julialang.org/u/quietlight)\
**Post date:** [August 19, 2023, 7:55am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/5 "2023-08-19T07:55:23Z")

</div>

Thanks for your reply. I will give it a go my tomorrow. Julia is an interesting language for sure.

Regards  
David

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 19, 2023, 6:18pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/6 "2023-08-19T18:18:12Z")

</div>

Yes `Tables.dictrowtable` is the intended way to handle this:

```julia
julia> DataFrame(Tables.dictrowtable(data))
4×3 DataFrame
 Row │ Number Name Address
     │ Int64 String String?
─────┼──────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 10 Jack Bog missing
   4 │ 5 Gid Hoo 120 Mut Street

```

(the issue is that your problem is unrelated with DataFrames.jl but is a consequence of Tables.jl design - there might be some more functionalities added to Tables.jl to make working with such data easier in the future)

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 19, 2023, 7:45pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/7 "2023-08-19T19:45:25Z")

</div>

Of the two unfulfilled expectations, of one (the one related to the function append!(…, cols=:union)) I got an idea of how it works.  
From the following example I understand that kwarg intervenes to combine internally “homogeneous” blocks (ie a vectors of named tuples with the same fields) but between the two o more blocks there may be fields not present in the others.  
I can’t figure out why the push!() function can’t work on a list of namedtuples instead

```julia
julia> df1=DataFrame(data[Not(3)])
3×3 DataFrame
 Row │ Number Name Address        
     │ Int64 String String
─────┼──────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 5 Gid Hoo 120 Mut Street

julia> append!(df1,data[3:3],cols=:union)
4×3 DataFrame
 Row │ Number Name Address        
     │ Int64 String String?
─────┼──────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 5 Gid Hoo 120 Mut Street
   4 │ 10 Jack Bog missing
#--------------------
julia> df1=DataFrame(data[1:1])
1×3 DataFrame
 Row │ Number Name Address    
     │ Int64 String String
─────┼──────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd

julia> append!(df1,data[2:4],cols=:union)
ERROR: type NamedTuple has no field Address
Stacktrace:
#-------------------
julia> df1=DataFrame(data[1:2])
2×3 DataFrame
 Row │ Number Name Address      
     │ Int64 String String
─────┼────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava

julia> append!(df1,data[3:4],cols=:union)
4×3 DataFrame
 Row │ Number Name Address      
     │ Int64 String String?
─────┼────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava
   3 │ 10 Jack Bog missing #???????
   4 │ 5 Gid Hoo missing

```

The behavior also seems to depend on the order in which the array of named tuples to be appended is prepared

```julia

julia> df1=DataFrame(data[1:2])
2×3 DataFrame
 Row │ Number Name Address      
     │ Int64 String String
─────┼────────────────────────────────
   1 │ 4 Abc Efgg 48 Mont Rd
   2 │ 6 Ruf Sly 19A Keke ava

julia> append!(df1,data[[4,3]],cols=:union)
ERROR: type NamedTuple has no field Address

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 19, 2023, 8:16pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/8 "2023-08-19T20:16:13Z")

</div>

> [@rocco\_sprmnt21](#):
>
> `push!(df,data...,cols=:union)`

This is not implemented as it strains compiler a lot, `foreach` should be used instead - just as you have proposed.

> `append!(df1,data[3:4],cols=:union)`

As I have commented - the reason why this fails is unrelated with DataFrames.jl. This is an issue with Tables.jl. The problem is that `data[3:4]` is not a valid Tables.jl table. That is why `Tables.dictrowtable` is currently required. However, in [Initializing Dataframe from vector of named tuples: missing values · Issue #3370 · JuliaData/DataFrames.jl · GitHub](https://github.com/JuliaData/DataFrames.jl/issues/3370) I proposed to add more support for non-homogenous tables in Tables.jl.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [August 19, 2023, 10:29pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/9 "2023-08-19T22:29:56Z")

</div>

> [@bkamins](#):
>
> The problem is that `data[3:4]` is not a valid Tables.jl table.

How to determine that? From reading the docs, it seems to fulfil the requirements for a row-based table.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 20, 2023, 7:29am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/10 "2023-08-20T07:29:52Z")

</div>

Isn’t the istable function checking if something is valid Tables.jl table?

```julia
julia> data = [(Number = 4, Name = "Abc Efgg", Address = "48 Mont Rd"),
       (Number = 6, Name = "Ruf Sly", Address = "19A Keke ava"),
       (Number = 10, Name = "Jack Bog"),
       (Number = 5, Name = "Gid Hoo", Address = "120 Mut Street")]
4-element Vector{NamedTuple}:
 (Number = 4, Name = "Abc Efgg", Address = "48 Mont Rd")
 (Number = 6, Name = "Ruf Sly", Address = "19A Keke ava")
 (Number = 10, Name = "Jack Bog")
 (Number = 5, Name = "Gid Hoo", Address = "120 Mut Street")

julia> Tables.istable(data)
true

julia> Tables.istable(data[3:4])
true

julia> Tables.istable(data[[4,3]])
true

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 8:09am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/11 "2023-08-20T08:09:09Z")

</div>

Technically you can check it if you try `Tables.columns`:

```julia
julia> Tables.columns(data)
ERROR: type NamedTuple has no field Address

```

In the documentation (of `dictrowtable`) you can read:

> For “schema-less” input tables, `dictrowtable` employs a “column unioning” behavior, as opposed to inferring the schema from the first row like `Tables.columns`.

So as you can read here normally the columns from a first row of data will be assumed to specify the columns of a table. That is why you get an error.

Also because of this if you have the following operation:

```julia
julia> DataFrame(data[[3, 1, 2, 4]])
4×2 DataFrame
 Row │ Number Name
     │ Int64 String
─────┼──────────────────
   1 │ 10 Jack Bog
   2 │ 4 Abc Efgg
   3 │ 6 Ruf Sly
   4 │ 5 Gid Hoo

```

it works and uses only 2 columns (from the 3rd row of the original table)

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 8:12am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/12 "2023-08-20T08:12:16Z")

</div>

> [@rocco\_sprmnt21](#):
>
> Isn’t the istable function checking if something is valid Tables.jl table?

I would not recommend using `Tables.istable` in practice (unfortunately). See its docs:

> Check if an object has specifically defined that it is a table. Note that not all valid tables will return true, since it’s possible to satisfy the Tables.jl interface at “run-time”

and

> It is recommended that for users implementing `MyType`, they define only `istable(::Type{MyType})`

so as you can see:

1. If `istable` returns `false` it does not mean anything
2. IF `istable` returns `true` it is typically determined on TYPE level (not instance level) and TYPE could have opted in to signal that it is a table, while the instance might violate some assumptions (this is the case of our `data` vector)

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 20, 2023, 8:18am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/13 "2023-08-20T08:18:28Z")

</div>

Would it be a bad idea (if it were possible) to transform (behind the scenes) the following expressions into the “good” one using dictrowtable?

```julia
   df=DataFrame()

    push!(df,data...,cols=:union)

    append!(df,data,cols=:union)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 8:23am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/14 "2023-08-20T08:23:52Z")

</div>

> [@rocco\_sprmnt21](#):
>
> `push!(df,data...,cols=:union)`

This is easily doable, by adding a following definition (simplified - simplified because we probably should not use recursion + we should handle all kwargs):

```julia
push!(df, d1, data...,cols) = push!(push!(df, d1, cols=cols), data..., cols=cols)

```

if you think it would be useful can you please open an issue?

* * *

For this:

> [@rocco\_sprmnt21](#):
>
> append!(df,data,cols=:union)

It cannot be fixed in DataFrames.jl. The reason is that it is Tables.jl that signals that `data` is not a valid table before even `append!` gets called. We would need something like `append!(df,Tables.colunion(data))` and add `Tables.colunion` to Tables.jl (`colunion` name is tentative).

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 20, 2023, 8:38am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/15 "2023-08-20T08:38:32Z")

</div>

to extend the use of the push function! I think it’s useful (if feasible without losing too much performance) as it’s a direct extension of how push!() works for “normal” arrays.  
I’ll open the issue right away. And I will be grateful if you make it easier for me by giving me the correct link of where to write it.

For the function append! I read from the documentation that

_Add the rows of df2 to the end of df. If the second argument table is not an AbstractDataFrame then it is converted using DataFrame(table, copycols=false) before being appended._

could you then (do I make it too easy? I don’t want to sound presumptuous in making suggestions to you. it’s just to understand something more) do DataFrame( dictrowtable(table), copycols=false) before being appended?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 8:56am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/16 "2023-08-20T08:56:17Z")

</div>

> `DataFrame( dictrowtable(table), copycols=false)`

This is very inefficient (computationally expensive). Therefore, what we now propose is:

- if you really need it you can add `Tables.dictrowtable` wrapper manually around `data` and things work
- to add another wrapper, I called it `Tables.colunion` tentatively, that would work like `Tables.dictrowtable` but would perform column unioning

I have opened the issue for the feature you asked for in [Add support for multiple positional arguments in push!/pushfirst!/append!/prepend! · Issue #3371 · JuliaData/DataFrames.jl · GitHub](https://github.com/JuliaData/DataFrames.jl/issues/3371).

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [August 20, 2023, 11:06am UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/17 "2023-08-20T11:06:52Z")

</div>

> [@bkamins](#):
>
> Technically you can check it if you try `Tables.columns`:

`data` seems to fulfill the requirements for a row-table: `rows(data)` is an iterable of AbstractRow-like objects. So, if `columns()` doesn’t work with it — either a bug in `columns`, or some requirements is missing in the docs (on implementing tables Interface).

From

> [@bkamins](#):
>
> For “schema-less” input tables, `dictrowtable` employs a “column unioning” behavior, as opposed to inferring the schema from the first row like `Tables.columns`.

it follows that “schema-less” tables are actually tables, their support is just not implemented in `columns`.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 1:56pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/18 "2023-08-20T13:56:04Z")

</div>

> [@aplavin](#):
>
> it follows that “schema-less” tables are actually tables, their support is just not implemented in `columns`.

If you feel the behavior should be changed, can you please open an issue in Tables.jl as probably @quinnj should comment on this since he maintains this package.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 20, 2023, 4:47pm UTC](https://discourse.julialang.org/t/construct-dataframe-from-uneven-named-tuples/102970/19 "2023-08-20T16:47:56Z")

</div>

> <https://github.com/JuliaData/DataFrames.jl/pull/3372>
>
> Fixes https://github.com/JuliaData/DataFrames.jl/issues/3371
> 
> Note that, for c…onsistency, I add support for \`cols\` in \`push!\`/\`pushfirst!\` when a collection without column names is pushed.
