# DataFrames/CSV: how to read vectors from \*.csv?

**URL:** <https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883>\
**Category:** General Usage\
**Created:** [March 24, 2021, 5:19pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883 "2021-03-24T17:19:15Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![mike\_k](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike_k/32/211864_2.png) [@mike\_k](https://discourse.julialang.org/u/mike_k)\
**Post date:** [March 24, 2021, 5:19pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/1 "2021-03-24T17:19:15Z")

</div>

Dear folks,  
unfortunately, I fail to read vectors from a \*.csv file. Here is an MWE:

```julia
using CSV, DataFrames

df = DataFrame( A=[[1,2,3], [4,5,6]] )
file = "my_test.csv"
CSV.write(file, df, delim=';')

new_df = DataFrame( CSV.File(file; delim=';' ) )

```

The problem is that column A in `new_df` of type `String`:

```julia
jjulia> new_df
2×1 DataFrame
 Row │ A         
     │ String    
─────┼───────────
   1 │ [1, 2, 3]
   2 │ [4, 5, 6]

```

I also tried:

```julia
CSV.File( file; types=Dict(:B => Vector{Int}) )

```

and

```julia
CSV.File( file; types=Dict(:B => Int[]) )

```

but it does not work. Do you have any ideas? Thanks in advance!

---

<div class="post-metadata">

**Author:** ![chris-b1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chris-b1/32/14165_2.png) [@chris-b1](https://discourse.julialang.org/u/chris-b1)\
**Post date:** [March 24, 2021, 5:49pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/2 "2021-03-24T17:49:47Z")

</div>

There currently isn’t a built-in way to do this with `CSV.jl` Take a look at this issue [comment](https://github.com/JuliaData/CSV.jl/issues/740#issuecomment-695128156) for a suggested approach, using a custom type.

Or perhaps more simply, you could parse the data after the fact

```julia
# read this inside out - strip off brackets, split by comma, then parse each sub-array to ints
julia> new_df.A_parsed = map(split.(strip.(new_df.A, Ref(['[', ']'])), ',')) do nums
           parse.(Int64, nums)
       end
2-element Array{Array{Int64,1},1}:
 [1, 2, 3]
 [4, 5, 6]

julia> new_df
2×2 DataFrame
 Row │ A A_parsed
     │ String Array…
─────┼──────────────────────
   1 │ [1, 2, 3] [1, 2, 3]
   2 │ [4, 5, 6] [4, 5, 6]

```

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [March 24, 2021, 7:22pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/3 "2021-03-24T19:22:44Z")

</div>

I do not know any CSV library that would do this for you. Is this use common in other languages?

---

<div class="post-metadata">

**Author:** ![mike\_k](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike_k/32/211864_2.png) [@mike\_k](https://discourse.julialang.org/u/mike_k)\
**Post date:** [March 24, 2021, 7:28pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/4 "2021-03-24T19:28:19Z")

</div>

I don’t know. But I think it would be a great feature. At least a certain CSV package should be able to “understand” the column types it wrote.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 24, 2021, 8:16pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/5 "2021-03-24T20:16:29Z")

</div>

Edit: comment removed, didn’t read OP closely enough.

Not sure I would generally expect this to work, however.

---

<div class="post-metadata">

**Author:** ![chris-b1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chris-b1/32/14165_2.png) [@chris-b1](https://discourse.julialang.org/u/chris-b1)\
**Post date:** [March 24, 2021, 9:49pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/6 "2021-03-24T21:49:31Z")

</div>

CSV is inevitably a lossy format - you might also consider `Arrow.jl` or some other binary format that can roundtrip this data.

```julia
using Arrow
Arrow.write("test.arrow", df)

julia> DataFrame(Arrow.Table("test.arrow"))
2×1 DataFrame
 Row │ A
     │ Array…
─────┼───────────
   1 │ [1, 2, 3]
   2 │ [4, 5, 6]

```

---

<div class="post-metadata">

**Author:** ![Mattriks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattriks/32/351_2.png) [@Mattriks](https://discourse.julialang.org/u/Mattriks)\
**Post date:** [March 25, 2021, 9:22am UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/7 "2021-03-25T09:22:30Z")

</div>

Another way is:

```julia
df2 = CSV.File(file, delim=';')
eval.(Meta.parse.(df2.A))

```

but don’t know if this is passé.

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [March 25, 2021, 12:38pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/8 "2021-03-25T12:38:28Z")

</div>

+1 to this. The kinds of places where it’s common to have columns of type array or map have all moved away from storing their data as CSV’s. At best, you can resurrect the old Hive conventions with explicit collection item delimiters: [CREATE HIVEFORMAT TABLE - Spark 3.3.0 Documentation](https://spark.apache.org/docs/latest/sql-ref-syntax-ddl-create-table-hiveformat.html)

---

<div class="post-metadata">

**Author:** ![cgeoga](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cgeoga/32/216186_2.png) [@cgeoga](https://discourse.julialang.org/u/cgeoga)\
**Post date:** [March 25, 2021, 2:21pm UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/9 "2021-03-25T14:21:17Z")

</div>

You could also just use `Serialization.serialize` and `Serialization.deserialize` from the standard library. From the docstrings it looks like they don’t promise compatibility across Julia versions, and you can’t open the serialized files in a text editor and look at what’s inside. But I think for use cases where you don’t want to use heavy machinery like JLD/HDF5/Arrow/etc, that’s a perfectly convenient way to store an object with a fancy type.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [March 26, 2021, 10:09am UTC](https://discourse.julialang.org/t/dataframes-csv-how-to-read-vectors-from-csv/57883/10 "2021-03-26T10:09:00Z")

</div>

In addition to the great suggestions above, you could also reformat as [“tidy” data](https://r4ds.had.co.nz/tidy-data.html), eg

```julia
6×2 DataFrame
│ Row │ a │ index │
│ │ Int64 │ Int64 │
├─────┼───────┼───────┤
│ 1 │ 1 │ 1 │
│ 2 │ 2 │ 2 │
│ 3 │ 3 │ 3 │
│ 4 │ 4 │ 1 │
│ 5 │ 5 │ 2 │
│ 6 │ 6 │ 3 │

```

which CSV should handle fine.
