# , and . in numbers inside CSV format

**URL:** <https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087>\
**Category:** Data\
**Tags:** question, csv\
**Created:** [February 5, 2023, 10:21am UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087 "2023-02-05T10:21:42Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![mpirke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mpirke/32/44837_2.png) [@mpirke](https://discourse.julialang.org/u/mpirke)\
**Post date:** [February 5, 2023, 10:21am UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/1 "2023-02-05T10:21:42Z")

</div>

Hey there!

I have to following problem. I want to read in data ( .csv ) which has the following structure:

```julia
data = """
col1;col2;col3;col4;col5
"05.02.2023";1000,01;2000,02;3000,03;12:00:00
"06.02.2023";4000,04;5000,05;6000,06;12:00:00
"""

```

This data can be easily read, by stating that ‘,’ is used as a decimal instead of ‘.’

```julia
file = CSV.File(IOBuffer(data); delim=';', decimal=',', dateformat="dd.mm.yyyy")

2-element CSV.File:
 CSV.Row: (col1 = Dates.Date("2023-02-05"), col2 = 1000.01, col3 = 2000.02, col4 = 3000.03, col5 = String15("12:00:00"))
 CSV.Row: (col1 = Dates.Date("2023-02-06"), col2 = 4000.04, col3 = 5000.05, col4 = 6000.06, col5 = String15("12:00:00"))

```

However my data looks like this:

```julia
data = """
col1;col2;col3;col4;col5
"05.02.2023";1.000,01;2.000,02;3.000,03;12:00:00
"06.02.2023";4.000,04;5.000,05;6.000,06;12:00:00
"""

```

There is an extra ‘.’ which separates digits of 1000, like it’s implemented in Julia with ‘\_’ underscore. So 1000,0 = 1.000,0 = 1\_000,0

Is there a way of handling such problems easily, without needing to do something like this":  
-read file line by line  
-do this where needed: replace(x, “.” =\> “”)  
-save file  
-then read in with CSV

Thanks for your help!

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [February 5, 2023, 11:50am UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/2 "2023-02-05T11:50:42Z")

</div>

It is on a to-do list [Ability to specify a thousands separator character · Issue #626 · JuliaData/CSV.jl · GitHub](https://github.com/JuliaData/CSV.jl/issues/626).

For now pass custom parser for this column. The parser can be from Parsers.jl:

```julia
julia> CSV.Parsers.parse(Float64, "1.234,56", CSV.Parsers.Options(groupmark='.', decimal=','))
1234.56

```

This will not be currently super convenient unfortunately.

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [February 5, 2023, 11:59am UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/3 "2023-02-05T11:59:36Z")

</div>

It seemed straightforward enough to make CSV.jl support this so I proposed [a PR](https://github.com/JuliaData/CSV.jl/pull/1072). I think it might land soon to the master branch.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 5, 2023, 1:47pm UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/4 "2023-02-05T13:47:58Z")

</div>

> [@mpirke](#):
>
> Is there a way of handling such problems easily, without needing to do something like this":  
> -read file line by line  
> -do this where needed: replace(x, “.” =\> “”)  
> -save file  
> -then read in with CSV

Unless the file is huge, you could just read it into a buffer, do the replacement, and read the CSV from an IOBuffer, rather than saving as another file. You can even do this in-place. e.g.

```julia
buf = filter!(!==(UInt8('.')), read("myfile.dat"))
file = CSV.File(IOBuffer(buf); delim=';', decimal=',')

```

---

<div class="post-metadata">

**Author:** ![mpirke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mpirke/32/44837_2.png) [@mpirke](https://discourse.julialang.org/u/mpirke)\
**Post date:** [February 5, 2023, 3:07pm UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/5 "2023-02-05T15:07:53Z")

</div>

Ah good to know. This would be a great feature!

And thanks for the advice with Parsers.jl, but as you said this is a little inconvenient.  
I’ll stick with the idea from @stevengj , this works ok, until such a feature is implemented in CSV.jl directly. 🙂

---

<div class="post-metadata">

**Author:** ![mpirke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mpirke/32/44837_2.png) [@mpirke](https://discourse.julialang.org/u/mpirke)\
**Post date:** [February 5, 2023, 3:09pm UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/6 "2023-02-05T15:09:40Z")

</div>

Thanks for the idea! This works well enough for me 🙂  
My files are not that big, the only disadvantage is that I also filter out the ‘.’ of the date, but this is not so bad.

Thanks!

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [February 5, 2023, 4:28pm UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/7 "2023-02-05T16:28:16Z")

</div>

> [@stevengj](#):
>
> `filter!(!==(UInt8('.')), read("myfile.dat"))`

This will work assuming you do not use `'.'` anywhere else in the file (@mpirke has mentioned that you encode dates as `05.02.2023` which would affect them)

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [February 5, 2023, 5:30pm UTC](https://discourse.julialang.org/t/and-in-numbers-inside-csv-format/94087/8 "2023-02-05T17:30:48Z")

</div>

A straightforward solution is to read the table into memory as-is, and then process remaining string columns.

```julia
julia> tbl = CSV.File(IOBuffer(data); delim=';', dateformat="dd.mm.yyyy") |> columntable

julia> tbl = map(tbl) do xs
           map(xs) do x
               x isa AbstractString && occursin(r"^[\d.,]+$", x) ?
                  parse(Float64, replace(x, '.' => "", ',' => '.')) :
                  x
           end
       end

```

This keeps datetimes and other columns that were successfully parsed by CSV.  
Kinda assumes that either all strings in a column match, or none do — otherwise you’ll get a mix of strings and floats.

The same code above works for both `columntable`s and `rowtable`s, as well as other table types like `StructArrays`.
