# Is there a way to read a DataFrame from file specifying the type of each column?

**URL:** <https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [November 1, 2024, 5:54pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146 "2024-11-01T17:54:40Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![world-peace](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Post date:** [November 1, 2024, 5:54pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/1 "2024-11-01T17:54:40Z")

</div>

Here is some code which reads a `DataFrame` from a (csv) file:

```julia
using DataFrames
using CSV

df = DataFrame(CSV.File(filename))

```

By default, I think the columns are `String` type. (Maybe it is inferred from the data values?)

… either way -

- Is there a way to specify the data types of each column when reading the data from file?

For this to work, obviously the column names need to be known in advance. In this case, that’s ok, because the `DataFrame` stores some data which has some constraints on its schema. There are a fixed set of column names, and the data type for each one is known in advance of reading the file from disk.

If this _isn’t_ possible - and I suspect it may not be - is there a way to _convert_ the columns using something equivalent to this python code:

```python
df = (
    df.astype(
        {
            'col_a': 'int',
            'col_b': 'string',
            'col_c': 'float',
        }
    )
)

```

---

<div class="post-metadata">

**Author:** ![eldee](https://avatars.discourse-cdn.com/v4/letter/e/b5a626/32.png) [@eldee](https://discourse.julialang.org/u/eldee)\
**Post date:** [November 1, 2024, 6:09pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/2 "2024-11-01T18:09:26Z")

</div>

Sure, this is possible, e.g. via `CSV.read`’s `types` argument.

> **Example**
>
> test.csv:
> 
> ```plaintext
> Col1,Col2,Col3
> 1,2.3,"sdf"
> 12,-0.213,"ds"
> 
> ```
> 
> ```julia-repl
> julia> using CSV, DataFrames
> 
> julia> CSV.read("test.csv", DataFrame)
> 2×3 DataFrame
> Row │ Col1 Col2 Col3
> │ Int64 Float64 String3
> ─────┼─────────────────────────
> 1 │ 1 2.3 sdf
> 2 │ 12 -0.213 ds
> 
> julia> CSV.read("test.csv", types=Dict(:Col1 => Float64), DataFrame)
> 2×3 DataFrame
> Row │ Col1 Col2 Col3
> │ Float64 Float64 String3
> ─────┼───────────────────────────
> 1 │ 1.0 2.3 sdf
> 2 │ 12.0 -0.213 ds
> 
> ```

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [November 1, 2024, 6:11pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/3 "2024-11-01T18:11:09Z")

</div>

Just read the documentation: [Reading · CSV.jl](https://csv.juliadata.org/stable/reading.html#types)

---

<div class="post-metadata">

**Author:** ![world-peace](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Post date:** [November 1, 2024, 6:12pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/4 "2024-11-01T18:12:09Z")

</div>

No way to do it using the `DataFrames` library? I _guess_ it makes sense that the type conversion is done at the time of reading the file though… Thanks for the tips… guess I was looking in the wrong place…

---

<div class="post-metadata">

**Author:** ![world-peace](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Post date:** [November 1, 2024, 6:19pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/5 "2024-11-01T18:19:37Z")

</div>

I think I have spoken too soon.

The problem with `CSV.read` is it requires a `filename` as the argument, whereas what I need is a `CSV.File`…

---

<div class="post-metadata">

**Author:** ![world-peace](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Post date:** [November 1, 2024, 6:21pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/6 "2024-11-01T18:21:31Z")

</div>

Tried this in the REPL… This seems to work…

```julia
CSV.File("example.txt", types=Dict(:col1=>Date,:col2=>String,:col3=>String,:col4=>Float64))

```

---

<div class="post-metadata">

**Author:** ![eldee](https://avatars.discourse-cdn.com/v4/letter/e/b5a626/32.png) [@eldee](https://discourse.julialang.org/u/eldee)\
**Post date:** [November 1, 2024, 6:36pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/7 "2024-11-01T18:36:00Z")

</div>

> [@world-peace](#):
>
> The problem with `CSV.read` is it requires a `filename` as the argument, whereas what I need is a `CSV.File`…

Why exactly do you want a `CSV.File`? If you want to end up with a `DataFrame` you can use `DataFrame` as `sink` in `CSV.read`.

> [@world-peace](#):
>
> is there a way to _convert_ the columns using something equivalent to this python code:

You could just manually convert after reading, using something like

```julia-repl
julia> df = CSV.read("test.csv", DataFrame) # initial DataFrame
2×3 DataFrame
 Row │ Col1 Col2 Col3
     │ Int64 Float64 String3
─────┼─────────────────────────
   1 │ 1 2.3 sdf
   2 │ 12 -0.213 ds

julia> for (col_symb, col_type) in zip((:Col1, :Col3), (Float32, String))
           df[!, col_symb] .= convert.(col_type, df[!, col_symb])
       end

julia> df
2×3 DataFrame
 Row │ Col1 Col2 Col3
     │ Float32 Float64 String
─────┼──────────────────────────
   1 │ 1.0 2.3 sdf
   2 │ 12.0 -0.213 ds

```

as in this topic

> [@DataFrames: convert column data type](https://discourse.julialang.org/t/dataframes-convert-column-data-type/35522/4):
>
> A more general way to do this is (assuming the column is called x) df[!,:x] = convert.(Bool,df[!,:x])

---

<div class="post-metadata">

**Author:** ![world-peace](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Post date:** [November 1, 2024, 7:37pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-read-a-dataframe-from-file-specifying-the-type-of-each-column/122146/8 "2024-11-01T19:37:19Z")

</div>

> [@eldee](#):
>
> Why exactly do you want a `CSV.File`? If you want to end up with a `DataFrame` you can use `DataFrame` as `sink` in `CSV.read`.

I have an interface which needs to take a file-like object, rather than a `DataFrame` directly.

> [@eldee](#):
>
> You could just manually convert after reading, using something like

That would work too, thanks for the suggestion.
