# Creating DataFrame from vector of rows, and fixing type

**URL:** <https://discourse.julialang.org/t/creating-dataframe-from-vector-of-rows-and-fixing-type/126591>\
**Category:** Data\
**Tags:** dataframes, type-inference\
**Created:** [March 5, 2025, 6:01pm UTC](https://discourse.julialang.org/t/creating-dataframe-from-vector-of-rows-and-fixing-type/126591 "2025-03-05T18:01:54Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![digital\_carver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/digital_carver/32/33818_2.png) [@digital\_carver](https://discourse.julialang.org/u/digital_carver)\
**Post date:** [March 5, 2025, 6:01pm UTC](https://discourse.julialang.org/t/creating-dataframe-from-vector-of-rows-and-fixing-type/126591/1 "2025-03-05T18:01:54Z")

</div>

I have some data in a `Vector{Vector{SubString{String}}}` that I wanted to turn into a DataFrame. While `DataFrame` constructor accepts a vector of vectors, it seems that it assumes the inner vectors are columns (which is sensible). To adapt it for my case, I landed on

```
DataFrame(stack(d; dims=1), columnnames)

```

(There doesn’t seem to be an `Iterators` version of `stack` afaict.)

1. Is this a reasonable way to go about this, or is there a better/more ergonomic way to handle this?

* * *

After I called the above, I ended up with a DataFrame where each column is of type `SubString`. Some of my columns are actually integers, and some are floats (and only one is an actual String).

1. What’s the best way to bring these columns to the appropriate type automatically? ~~I vaguely remember a trick using the `identity` function, but not exactly how to use it, nor do I know if that’s the recommended path here.~~ Actually, I believe that trick was for when `missing`s are removed or the type otherwise needs narrowing, not for when it needs to be actually changed. So I guess the path here is to just parse the columns into the right type and replace them individually?

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 6, 2025, 6:25pm UTC](https://discourse.julialang.org/t/creating-dataframe-from-vector-of-rows-and-fixing-type/126591/2 "2025-03-06T18:25:08Z")

</div>

Waiting for a more specific and julianic solution  
I want to imagine that the vector of vectors was obtained by splitting a text.  
Then operating inversely …

```julia
using DataFrames,CSV
vov=[["a","1","2."],["b","3","4."],["c","5","6."]]
str=join(join.(vov,','),'\n')
write("df.txt",str)
CSV.read("df.txt", DataFrame,header=false)

```

```julia
julia> CSV.read("df.txt", DataFrame,header=false)
3×3 DataFrame
 Row │ Column1 Column2 Column3 
     │ String1 Int64 Float64
─────┼───────────────────────────
   1 │ a 1 2.0
   2 │ b 3 4.0
   3 │ c 5 6.0

```

to save disk space :).  
But I don’t think it’s a good habit

```julia
julia> io=IOBuffer()
IOBuffer(data=UInt8[...], readable=true, writable=true, seekable=true, append=false, size=0, maxsize=Inf, ptr=1, mark=-1)

julia> write(io,str)
20

julia> CSV.read(io.data[1:io.size], DataFrame,header=false)
3×3 DataFrame
 Row │ Column1 Column2 Column3 
     │ String1 Int64 Float64
─────┼───────────────────────────
   1 │ a 1 2.0
   2 │ b 3 4.0
   3 │ c 5 6.0

```

---

<div class="post-metadata">

**Author:** ![bertschi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bertschi/32/33462_2.png) [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Post date:** [March 6, 2025, 7:38pm UTC](https://discourse.julialang.org/t/creating-dataframe-from-vector-of-rows-and-fixing-type/126591/3 "2025-03-06T19:38:17Z")

</div>

If you are starting from strings, you will need to parse the data somehow. Here, is my take

```julia-repl
julia> stuff = [["1", view("one", 1:3)], ["2", "two"], ["3", "three"]]
3-element Vector{Vector{AbstractString}}:
 ["1", "one"]
 ["2", "two"]
 ["3", "three"]

julia> spec = (x = Base.Fix1(tryparse, Int64), y = identity);

julia> Iterators.map(x -> NamedTuple{keys(spec)}(x .|> values(spec)), stuff) |> DataFrame
3×2 DataFrame
 Row │ x y         
     │ Int64 Abstract… 
─────┼──────────────────
   1 │ 1 one
   2 │ 2 two
   3 │ 3 three

```

Not claiming that this is efficient.
