# Save Dataframe in file and read it again

**URL:** <https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336>\
**Category:** General Usage\
**Tags:** question\
**Created:** [May 28, 2020, 1:22pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336 "2020-05-28T13:22:58Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![FrootLoops](https://avatars.discourse-cdn.com/v4/letter/f/c4cdca/32.png) [@FrootLoops](https://discourse.julialang.org/u/FrootLoops)\
**Post date:** [May 28, 2020, 1:22pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336/1 "2020-05-28T13:22:58Z")

</div>

Hello,

right now I am trying to save a Dataframe into a file and read it again. For this I used the CSV package. My dataframe has two columns where each entry is a vector. Here is an small example:

```julia
using DataFrames
using CSV

x = Vector{Array{Float64, 1}}()
y = Vector{Array{Float64, 1}}()
push!(x, [1,2,3])
push!(x, [4,5,6])
push!(y, [100,101,102])
push!(y, [103,104,105])

df = DataFrame(x=x, y=y)
CSV.write("export_df.csv", df)

test = CSV.read("export_df.csv")
println(typeof(df["x"][1]))
println(typeof(test["x"][1]))

```

The first output returns an Array{Float64,1} type, while type of test[“x”][1] returns a string. Can someone give me an advise how to fix this, such that I get an Array again?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 28, 2020, 1:30pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336/2 "2020-05-28T13:30:42Z")

</div>

If your columns are Vector of Arrays then I think you would need something like JLSO.jl, JLD.jl or JLD2.jl or just serialize it

```julia
using DataFrames
x = Vector{Array{Float64, 1}}()
y = Vector{Array{Float64, 1}}()
push!(x, [1,2,3])
push!(x, [4,5,6])
push!(y, [100,101,102])
push!(y, [103,104,105])

df = DataFrame(x=x, y=y)

using Serialization
serialize("ok.jls", df)
df2 = deserialize("ok.jls")

df2 == df # true

```

But if you actually want to store just `Vector{Float64}` etc then you can try the [JDF.jl package](https://github.com/xiaodaigh/JDF.jl)? Other options are Parquet.jl and Feather.jl.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [May 28, 2020, 1:32pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336/3 "2020-05-28T13:32:55Z")

</div>

cc @quinnj on this. Reading through the docs CSV.File (which will replace CSV.read eventually), it looks like the `type` keyword you would use to specify types only works for small concrete types like `Float64`, `Int` etc.

I also think it’s tough to parse the results from `test` into an array.

My recommendation is to use DataFrame’s `flatten` function to save the data in “long” format and then use `combine` to put the result back into vectors after you read it in.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 28, 2020, 1:35pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336/4 "2020-05-28T13:35:09Z")

</div>

> [@pdeffebach](#):
>
> CSV.File

Presumably, the CSV format isn’t really designed to handle this. Serialization is a better option here IMO.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [May 28, 2020, 1:45pm UTC](https://discourse.julialang.org/t/save-dataframe-in-file-and-read-it-again/40336/5 "2020-05-28T13:45:22Z")

</div>

I would agree. The `flatten` option is pretty tough to implement and hard to generalize.

```julia
ulia> function save_to_csv(df)
           df.rownum = 1:nrow(df)
           long = stack(df, [:x, :y])
           longer = flatten(long, :value)
           CSV.write("out.csv", longer)
           return longer
       end

julia> function read_from_csv(file)
       df = CSV.File(file) |> DataFrame
       combine(groupby(df, "rownum")) do sdf
           (x = [sdf.value[sdf.variable .== "x"]], 
            y = [sdf.value[sdf.variable .== "y"]])
       end
       end

```
