# Construct Julia Dataframe from row data

**URL:** <https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219>\
**Category:** New to Julia\
**Tags:** question, dataframes, data\_structures\
**Created:** [March 19, 2020, 7:06pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219 "2020-03-19T19:06:29Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![nnn](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@nnn](https://discourse.julialang.org/u/nnn)\
**Post date:** [March 19, 2020, 7:06pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/1 "2020-03-19T19:06:29Z")

</div>

When I have data in columns, I can construct a DataFrame like this:

```julia
using Dataframes
nameData = ["Alice", "Bob"]
surnameData = ["Smith", "Jones"]
df = DataFrame(name = nameData, surname = surnameData)

```

However I have data in rows:

```julia
row1 = ["Alice", "Smith"]
row2 = ["Bob", "Jones"]
df = DataFrame(???)

```

How can I make a DataFrame from it?

To generalize a question I have a large amount of row-aligned data in memory but outside of Julia and I need to construct a Julia DataFrame from it. What’s the fastest way to do that? Rows are always the same size but can contain different data types. Same column is always the same datatype for all rows.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 19, 2020, 7:31pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/2 "2020-03-19T19:31:31Z")

</div>

Two options

1. Initialize a data frame with no rows like `DataFrame(name = String[], surname = String[]`, then do

```julia
for row in rows # your collection of rows
    push!(df, row)
end

```

1. Write a small helper function that turns each `row` into a named tuple

```julia
function make_better_row(vec)
    (name = vec[1], surname = vec[2])
end

```

with this, you don’t have to worry about the types of vectors when you initialize the dataframe. you can do

```julia
df = DataFrame()
for row in rows
    push!(df, make_better_row(row))
end

```

This is all assuming your rows come to you where each row is a vector. But you get the idea: push a named tuple to an empty data frame (i.e. `df = DataFrame())`, but if you are pushing a vector or an ordinary tuple, initialize the dataframe’s columns first.

---

<div class="post-metadata">

**Author:** ![lbilli](https://avatars.discourse-cdn.com/v4/letter/l/59ef9b/32.png) [@lbilli](https://discourse.julialang.org/u/lbilli)\
**Post date:** [March 19, 2020, 7:40pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/3 "2020-03-19T19:40:01Z")

</div>

If you can put rows into a vector of named tuples:

```julia
julia> row1 = (name="Alice", age=20)
julia> row2 = (name="Bob", age=30)

julia> DataFrame([row1, row2])
2×2 DataFrame
│ Row │ name │ age │
│ │ String │ Int64 │
├─────┼────────┼───────┤
│ 1 │ Alice │ 20 │
│ 2 │ Bob │ 30 │

```

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [March 19, 2020, 8:35pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/4 "2020-03-19T20:35:09Z")

</div>

Is this of any use here

> [@\[ANN\] DataFrameDBs.jl](https://discourse.julialang.org/t/ann-dataframedbs-jl/35718/5):
>
> I don’t work with kdb+, DataFrameDBs is primarily inspired by [ClickHouse](https://clickhouse.tech/) (and my own experience on developing columnar DB with c++). Performance of DataFrameDBs competes with ClickHouse when ClickHouse runs on a single thread.

---

<div class="post-metadata">

**Author:** ![waralex](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/waralex/32/13415_2.png) [@waralex](https://discourse.julialang.org/u/waralex)\
**Post date:** [March 19, 2020, 9:28pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/5 "2020-03-19T21:28:47Z")

</div>

And what is the problem of converting rows to columns?  
You can push rows to DataFrame as @pdeffebach describes, But this may not be very effective, because it will lead to several resizing of columns vectors.  
I would do something like that (I assume that rows is an array of arrays):

```julia
julia> columns = Vector{Vector}(undef, 0)
julia> for (i,r) in enumerate(first(rows))
           column = Vector{typeof(r)}(undef, length(rows))
           column .= getindex.(rows, i)
           push!(columns, column)
       end
julia> df = DataFrame(columns, column_names, copycols=false)

```

` column_names` must be Vector{Symbol}, not Vector{String}.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 19, 2020, 9:40pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/6 "2020-03-19T21:40:01Z")

</div>

One thing to remember is that Julia is very good at starting something from nothing and `push!`ing them, julia doubles the memory used by a vector each time it runs out of space. So it’s performance shouldn’t be an issue.

But this also works.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 19, 2020, 9:46pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/7 "2020-03-19T21:46:35Z")

</div>

Since this `push!` is pushing to DataFrame, it is possible that there can be other implementation issues. At least there is a recommendation in DataFrames manual to use columns approach: [Getting Started · DataFrames.jl](https://juliadata.github.io/DataFrames.jl/stable/man/getting_started/#Constructing-Row-by-Row-1)

---

<div class="post-metadata">

**Author:** ![waralex](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/waralex/32/13415_2.png) [@waralex](https://discourse.julialang.org/u/waralex)\
**Post date:** [March 19, 2020, 10:09pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/8 "2020-03-19T22:09:48Z")

</div>

> [@pdeffebach](#):
>
> So it’s performance shouldn’t be an issue.

It depends on the number of rows

> [@pdeffebach](#):
>
> julia doubles the memory used by a vector each time it runs out of space

It was this process that I meant by “several resizes”. This is the usual behavior for arrays in most languages. It is not as fast as it seems. Each time you need to allocate a new memory block and copy data to it . Both are not instant operations at all. Assume that you already have an array of 2 ^ 30 bytes and you need to allocate 2 GB and copy 1 GB at next push!.. So if you know the size of the array, it’s a good idea to preallocate it. `sizehint!` designed just for this.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [March 19, 2020, 10:18pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/9 "2020-03-19T22:18:22Z")

</div>

Yet another way:

```julia
rows = [["Alice", "Smith"], ["Bob", "Jones"]]
DataFrame((name = r[1], surname = r[2]) for r in rows)

```

---

<div class="post-metadata">

**Author:** ![nnn](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@nnn](https://discourse.julialang.org/u/nnn)\
**Post date:** [March 20, 2020, 4:21pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/10 "2020-03-20T16:21:02Z")

</div>

Thank you for your suggestions. I’m checking them out.

Meanwhile I have continued my search and found out that CSV package also produces DataFrames and (presumably) engineered for row-aligned data. It also accepts IOBuffer as an input. However I encountered a problem.  
This code works and produces 2x2 DataFrame:

```julia
using CSV
rowData = "Alice,20\nBob,30"
CSV.read(IOBuffer(rowData), header=["Name","Age"])

```

But this code does not:

```julia
using CSV
rowData = "Alice,20\nBob,30"
io = IOBuffer()
write(io, rowData)
CSV.read(io, header=["Name","Age"])

```

Or rather it produces 0x2 DataFrame without rows. Why?

As I can potentially have a bigger dataset, I’d prefer to fill IOBuffer gradually, rather than once at initialization.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [March 21, 2020, 12:36am UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/11 "2020-03-21T00:36:44Z")

</div>

> [@nnn](#):
>
> Or rather it produces 0x2 DataFrame without rows. Why?

I’m guessing because the io pointer is at the end of the buffer, so when the reader looks, it sees the end immediately and closes. Try `seek(io, 0)` before calling `CSV.read()` to move the pointer back to the beginning

---

<div class="post-metadata">

**Author:** ![waralex](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/waralex/32/13415_2.png) [@waralex](https://discourse.julialang.org/u/waralex)\
**Post date:** [March 21, 2020, 8:34pm UTC](https://discourse.julialang.org/t/construct-julia-dataframe-from-row-data/36219/12 "2020-03-21T20:34:05Z")

</div>

```julia
julia> using CSV
julia> rowData = "Alice,20\nBob,30"
"Alice,20\nBob,30"
julia> io = PipeBuffer()
IOBuffer(data=UInt8[...], readable=true, writable=true, seekable=false, append=true, size=0, maxsize=Inf, ptr=1, mark=-1)
julia> write(io, rowData)
15
julia> CSV.read(io, header=["Name","Age"])
2×2 DataFrames.DataFrame
│ Row │ Name │ Age │
│ │ String │ Int64 │
├─────┼────────┼───────┤
│ 1 │ Alice │ 20 │
│ 2 │ Bob │ 30 │

```
