# Convert a percentage string (e.g. "99.8%") to Float64 when extracting data from CSV

**URL:** <https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278>\
**Category:** New to Julia\
**Created:** [September 8, 2020, 4:50pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278 "2020-09-08T16:50:15Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![alexlowth](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@alexlowth](https://discourse.julialang.org/u/alexlowth)\
**Post date:** [September 8, 2020, 4:50pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/1 "2020-09-08T16:50:15Z")

</div>

Hi,

Does anyone know if there is a possible way to convert a percentage from a string to a Float64 when extracting data from a CSV. Or is the only way to convert the data to decimal format in excel before pulling through to Julia?

I have tried using the parse function, but that doesn’t work.

As a simplification, I have replaced the CSV extract with the raw string to help visualise:

parse(Float64, “99.8%”) which produces the error ArgumentError: cannot parse “99.8%” as Float64\*

Thanks,  
Alex

---

<div class="post-metadata">

**Author:** ![anon37204545](https://avatars.discourse-cdn.com/v4/letter/a/439d5e/32.png) [@anon37204545](https://discourse.julialang.org/u/anon37204545)\
**Post date:** [September 8, 2020, 4:55pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/2 "2020-09-08T16:55:04Z")

</div>

`99.8%` is not an acceptable float. What you want is to to take `str[1:end-1]` (all characters but the last one) and then multiply it by 0.01.

```julia
julia> parse(Float64, "99.8%"[1:end-1]) * 0.01
0.998

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [September 8, 2020, 4:55pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/3 "2020-09-08T16:55:13Z")

</div>

You have to get rid of the “%”:

```julia
julia> parse(Float64,replace("99.8%","%"=>""))
99.8

```

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [September 8, 2020, 7:05pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/4 "2020-09-08T19:05:38Z")

</div>

No Julia thread would be complete without discussing performance. I tried another version of this using RegEx and it’s _slower_ than @lmiq’s version 😄:

```julia
using BenchmarkTools

values = ["$(rand(0.0:0.01:100.0))%" for _ in 1:10_000]

julia> @btime [parse(Float64, match(r"[-+]?[0-9]*\.?[0-9]+", val).match) for val in $values]
  6.252 ms (30002 allocations: 2.21 MiB)
10000-element Array{Float64,1}:
 45.26
 30.06
 60.07
 12.7
 59.45
 29.99
 80.86
 83.66
 77.91
  ⋮

julia> @btime [parse(Float64, replace(val,"%" => "")) for val in $values]
  5.855 ms (40002 allocations: 2.21 MiB)
10000-element Array{Float64,1}:
 45.26
 30.06
 60.07
 12.7
 59.45
 29.99
 80.86
 83.66
 77.91
  ⋮

```

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [September 8, 2020, 8:01pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/5 "2020-09-08T20:01:31Z")

</div>

> [@anon37204545](#):
>
> `99.8%` is not an acceptable float. What you want is to to take `str[1:end-1]` (all characters but the last one) and then multiply it by 0.01.
> 
> ```julia
> julia> parse(Float64, "99.8%"[1:end-1]) * 0.01
> 0.998
> 
> ```

Note that this does something different than the other two answers, rendering the percent as a fraction (eg 95% becomes 0.95). Also, in general, I would get out of the habit of indexing into strings where possible. It’s convenient, but can cause unexpected results of you have characters that use multiple code points. In this case, it’s probably fine.

---

<div class="post-metadata">

**Author:** ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Post date:** [September 8, 2020, 8:24pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/6 "2020-09-08T20:24:51Z")

</div>

Shouldn’t percentages directly be “understood” by the parse function?

It has not been implemented because lack of time/interest, or some clever reasons?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 8, 2020, 8:28pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/7 "2020-09-08T20:28:26Z")

</div>

I checked R and Stata and neither have those implemented, so it’s not the norm.

I think there is room for a `parse` function that gets rid of all non-digits and non `.` or `,`

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [September 8, 2020, 8:42pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/8 "2020-09-08T20:42:25Z")

</div>

Trying to reimplement Excel in the CSV reader seems like a bad road to go down. Should CSV.jl also start interpreting strings like SEPT1 as dates? If we lean in hard maybe we can get some [genes renamed](https://www.bionews.org.uk/page_151346).

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [September 8, 2020, 8:47pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/9 "2020-09-08T20:47:24Z")

</div>

As a more general comment, I really don’t understand trying to cram all this intelligence and performance into CSV readers. CSV is an awful format that requires parsing and guessing. If you get data in CSV format, you should parse it exactly once and then save it in a sane format that can just be loaded without all that nonsense. If there are columns that use `99.8%` as a number format, load them as strings, parse those strings and convert to numbers and then save the result in a reasonable format that doesn’t have all these problems. If you are repeatedly using a CSV parser on the same data, you have only yourself to blame if that takes longer than you would like it to.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [September 8, 2020, 11:11pm UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/11 "2020-09-08T23:11:11Z")

</div>

insane, does it mean:

1. people who deal with genes don’t know how to use Excel (to set columns to “text”)
2. people USE EXCEL TO do genetic stuff?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 9, 2020, 12:44am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/12 "2020-09-09T00:44:26Z")

</div>

> [@StefanKarpinski](#):
>
> As a more general comment, I really don’t understand trying to cram all this intelligence and performance into CSV readers. CSV is an awful format that requires parsing and guessing. If you get data in CSV format, you should parse it exactly once and then save it in a sane format that can just be loaded without all that nonsense. If there are columns that use `99.8%` as a number format, load them as strings, parse those strings and convert to numbers and then save the result in a reasonable format that doesn’t have all these problems. If you are repeatedly using a CSV parser on the same data, you have only yourself to blame if that takes longer than you would like it to.

For me it’s more valuable to have a separate util outside of CSV that can DETECT these issues and suggest improvements. I mean xx.x% are easy patterns to support. Of course, these utils can work with CSV but making it easy to transform the column at read time, but I’d say keep CSV.jl lean .

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [September 9, 2020, 6:37am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/13 "2020-09-09T06:37:33Z")

</div>

Maybe one could write a small package to implement a `Percent` type, and make it possible to do

```julia
parse(Percent{Float64}, "99.5%")

```

Parsing CSV would then work as long as the user specifies which columns should be parsed as this type. A missing “%” at the end of the string would then be a parse error.

(The division by 100 should not happen until converting to another numeric type, since 1//100 is not exactly representable in floating-point.)

---

<div class="post-metadata">

**Author:** ![anon37204545](https://avatars.discourse-cdn.com/v4/letter/a/439d5e/32.png) [@anon37204545](https://discourse.julialang.org/u/anon37204545)\
**Post date:** [September 9, 2020, 7:23am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/14 "2020-09-09T07:23:27Z")

</div>

> [@Per](#):
>
> Maybe one could write a small package to implement a `Percent` type, and make it possible to do

I tried this once, but I got stuck at promotion and conversion. How should such type act in respect to other `Number`s?

---

<div class="post-metadata">

**Author:** ![alexlowth](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@alexlowth](https://discourse.julialang.org/u/alexlowth)\
**Post date:** [September 9, 2020, 8:28am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/15 "2020-09-09T08:28:08Z")

</div>

You say to parse a string that contains “99.8%” but the point of the post was that this exact thing can’t be parsed? Sorry if misunderstood

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 9, 2020, 8:37am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/16 "2020-09-09T08:37:54Z")

</div>

> [@jling](#):
>
> people USE EXCEL TO do genetic stuff?

That’s just the tip of the iceberg. The more you learn about what people actually use Excel for, the more you become convinced that the current technological human civilization and specifically scientific research as such is inevitably doomed.

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [September 9, 2020, 8:58am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/17 "2020-09-09T08:58:09Z")

</div>

Promotion is a bit tricky when introducing new `Number` types. I would start with something simple like

```julia
promote_rule(::Type{Percent{S}}, ::Type{T}) where {S,T <: Number} = promote_type(S,T)

```

but there would likely be problems with ambiguities that would need to be ironed out.

It would also be nice to have

```julia
promote_rule(::Type{Percent{S}}, ::Type{Rational{T}}) where {S <: Integer, T <: Integer} = Rational{promote_type(S,T)}

```

so that 1% + 1//2 resulted in exactly 51//100 instead of approximately 0.51000000000000001, but this is probably being overly pedantic.

---

<div class="post-metadata">

**Author:** ![alexlowth](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@alexlowth](https://discourse.julialang.org/u/alexlowth)\
**Post date:** [September 9, 2020, 9:09am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/18 "2020-09-09T09:09:14Z")

</div>

Hi Zlatan, thanks for the solution. I guess the next question would be do you know how to do this on a vector of percentage strings? My understanding is that this would work for a single string

E.g: a1\_rfr\_disc  
Vector{Union{Missing, String}} with 10 elements  
“99.8%”  
“98.8%”  
“97.1%”  
“94.6%”  
“91.7%”  
“88.5%”  
“85.3%”  
“82.1%”  
“78.9%”  
“75.9%”

---

<div class="post-metadata">

**Author:** ![anon37204545](https://avatars.discourse-cdn.com/v4/letter/a/439d5e/32.png) [@anon37204545](https://discourse.julialang.org/u/anon37204545)\
**Post date:** [September 9, 2020, 9:11am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/19 "2020-09-09T09:11:30Z")

</div>

A clear solution would be to make a method which does what I described:

```julia
julia> percent(str) = parse(Float64, str[1:end-1]) * 0.01
percent (generic function with 1 method)

```

and then broadcast it (using `.`) to a vector:

```julia
julia> percent.(["99.8%", "99.7%"])
2-element Array{Float64,1}:
 0.998
 0.997

```

---

<div class="post-metadata">

**Author:** ![alexlowth](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@alexlowth](https://discourse.julialang.org/u/alexlowth)\
**Post date:** [September 9, 2020, 9:13am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/20 "2020-09-09T09:13:13Z")

</div>

Ah yes, I didn’t think of that but makes sense. Thanks

---

<div class="post-metadata">

**Author:** ![bernhard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bernhard/32/2619_2.png) [@bernhard](https://discourse.julialang.org/u/bernhard)\
**Post date:** [September 9, 2020, 10:37am UTC](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278/21 "2020-09-09T10:37:30Z")

</div>

Actually it mainly means that Excel does some very strange stuff to data, which can be difficult to prevent.  
Eg when opening a CSV file Excel will automatically format the data and there is nothing you can do against it (iirc).

[Next page](https://discourse.julialang.org/t/convert-a-percentage-string-e-g-99-8-to-float64-when-extracting-data-from-csv/46278.md?page=2)
