# Outputing/Inputing vectors in DataFrames

**URL:** https://discourse.julialang.org/t/outputing-inputing-vectors-in-dataframes/93099
**Category:** General Usage
**Tags:** dataframes, csv
**Created:** [January 17, 2023, 2:47pm UTC](https://discourse.julialang.org/t/outputing-inputing-vectors-in-dataframes/93099 "2023-01-17T14:47:26Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![mb96](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mb96/32/45786_2.png) [@mb96](https://discourse.julialang.org/u/mb96)
#### Post date: [January 17, 2023, 2:47pm UTC](https://discourse.julialang.org/t/outputing-inputing-vectors-in-dataframes/93099/1 "2023-01-17T14:47:26Z")

</div>

Hi,

After some calculations I define a DataFrame to store my results and output them as a CSV file.

 ![Screenshot 2023-01-17 at 9.41.36 AM](https://global.discourse-cdn.com/julialang/original/3X/1/8/18665d49b4e9c1066b4691ad349e42561c792f09.png)

The cutoffs columns are vectors. After outputting them as a CSV file I would like to read the file again an access the vectors as vectors but they are stored as Strings.

 ![Screenshot 2023-01-17 at 9.47.00 AM](https://global.discourse-cdn.com/julialang/original/3X/8/4/84e5971c37ff38474955829db2909800c1c1d35c.png)

Is there a way to tell DataFrames to store them as vectors such that when I read them I can access them as vectors? Or, is there a way to read in the output as string and convert them into vectors?

Thanks a lot,  
M.

---

<div class="post-metadata">

### Author: ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)
#### Post date: [January 17, 2023, 3:02pm UTC](https://discourse.julialang.org/t/outputing-inputing-vectors-in-dataframes/93099/2 "2023-01-17T15:02:51Z")

</div>

This is a very old question to which the answer is no: if you want to store additional structure with your data, don’t use csv.

See e.g. this prior discussion:

> [@DataFrames/CSV: how to read vectors from \*.csv?](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883):
>
> Dear folks, unfortunately, I fail to read vectors from a \*.csv file. Here is an MWE: using CSV, DataFrames df = DataFrame( A=[[1,2,3], [4,5,6]] ) file = "my\_test.csv" CSV.write(file, df, delim=';') new\_df = DataFrame( CSV.File(file; delim=';' ) ) The problem is that column A in new\_df of type String: jjulia\> new\_df 2×1 DataFrame Row │ A │ String ─────┼─────────── 1 │ [1, 2, 3] 2 │ [4, 5, 6] I also tried: CSV.File( file; types=Dict(:B =\> Vector{Int}) ) and CSV.F…

---

<div class="post-metadata">

### Author: ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)
#### Post date: [January 17, 2023, 3:03pm UTC](https://discourse.julialang.org/t/outputing-inputing-vectors-in-dataframes/93099/3 "2023-01-17T15:03:57Z")

</div>

If you want to use CSV as a storage format then I would not advise you to try what you do (however, you could save this data e.g. as Parquet2.jl file or Arrow.jl file).

In CSV what you should do is the following (as CSV does not support storing vectors of vectors):

```julia
julia> df = DataFrame(x=rand(3), y=[rand(4) for i in 1:3])
3×2 DataFrame
 Row │ x y
     │ Float64 Array…
─────┼─────────────────────────────────────────────
   1 │ 0.493174 [0.839034, 0.0397448, 0.766962, …
   2 │ 0.146369 [0.398974, 0.416861, 0.444297, 0…
   3 │ 0.729989 [0.33982, 0.629735, 0.313111, 0.…

julia> df.id = axes(df, 1) # add id column
Base.OneTo(3)

julia> df_to_save = flatten(df, :y) # flatten the data so that you can write it
12×3 DataFrame
 Row │ x y id
     │ Float64 Float64 Int64
─────┼────────────────────────────
   1 │ 0.493174 0.839034 1
   2 │ 0.493174 0.0397448 1
   3 │ 0.493174 0.766962 1
   4 │ 0.493174 0.955614 1
   5 │ 0.146369 0.398974 2
   6 │ 0.146369 0.416861 2
   7 │ 0.146369 0.444297 2
   8 │ 0.146369 0.765362 2
   9 │ 0.729989 0.33982 3
  10 │ 0.729989 0.629735 3
  11 │ 0.729989 0.313111 3
  12 │ 0.729989 0.220889 3

julia> CSV.write("test.csv", df_to_save) # save it
"test.csv"

julia> df_loaded = CSV.read("test.csv", DataFrame) # load it
12×3 DataFrame
 Row │ x y id
     │ Float64 Float64 Int64
─────┼────────────────────────────
   1 │ 0.493174 0.839034 1
   2 │ 0.493174 0.0397448 1
   3 │ 0.493174 0.766962 1
   4 │ 0.493174 0.955614 1
   5 │ 0.146369 0.398974 2
   6 │ 0.146369 0.416861 2
   7 │ 0.146369 0.444297 2
   8 │ 0.146369 0.765362 2
   9 │ 0.729989 0.33982 3
  10 │ 0.729989 0.629735 3
  11 │ 0.729989 0.313111 3
  12 │ 0.729989 0.220889 3

julia> combine(groupby(df_loaded, :id), :x => first, :y => Ref∘copy, keepkeys=false) # reverse the flattening - you might not need it, but in case you do this is how you can do it
3×2 DataFrame
 Row │ x_first y_Ref_copy
     │ Float64 Array…
─────┼─────────────────────────────────────────────
   1 │ 0.493174 [0.839034, 0.0397448, 0.766962, …
   2 │ 0.146369 [0.398974, 0.416861, 0.444297, 0…
   3 │ 0.729989 [0.33982, 0.629735, 0.313111, 0.…

```
