# Extracting a float from a string

**URL:** <https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126>\
**Category:** New to Julia\
**Tags:** strings, regex\
**Created:** [July 15, 2020, 12:45pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126 "2020-07-15T12:45:01Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![erlebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erlebach/32/12973_2.png) [@erlebach](https://discourse.julialang.org/u/erlebach)\
**Post date:** [July 15, 2020, 12:45pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/1 "2020-07-15T12:45:01Z")

</div>

I would like to extract a float from a string of the form  
`strg = file_t=0.234`. In Python, I would use a package called `scanf` with similar functionality to what is offered in C++. I might also use regular expressions, but that is the last resort. What solutions does Julia offer besides regular expressions? If any. Thanks.

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2020, 12:49pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/2 "2020-07-15T12:49:17Z")

</div>

```julia
strg = "file_t=0.234"
parse(Float64,split(strg,'=')[2])

```

---

<div class="post-metadata">

**Author:** ![erlebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erlebach/32/12973_2.png) [@erlebach](https://discourse.julialang.org/u/erlebach)\
**Post date:** [July 15, 2020, 12:56pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/3 "2020-07-15T12:56:26Z")

</div>

Thank you! I knew I’d use parse, and split, but never thought of using the [2] selector.

---

<div class="post-metadata">

**Author:** ![erlebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erlebach/32/12973_2.png) [@erlebach](https://discourse.julialang.org/u/erlebach)\
**Post date:** [July 15, 2020, 12:59pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/4 "2020-07-15T12:59:20Z")

</div>

Oops. My minimal example was too minimal. How about extracting the flow from the string: `"time t=0.234_ext.bson"`?  
I thought that `parse` would extract the first float it finds, similar to `scanf` in C++, rather than require that the string parsed was a float.  
Thanks.

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2020, 1:04pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/5 "2020-07-15T13:04:01Z")

</div>

If we stick with split and a fixed format we could do:

```julia
julia> strg="time t=0.234_ext.bson"
"time t=0.234_ext.bson"
julia> parse(Float64,split(split(strg,'=')[2],'_')[1])
0.234

```

If the format isn’t as fixed I would do a regular expression. Why do you exclude them?

For the first example there is another one (besides multiple others):

```julia
julia> strg = "file_t=0.234"
"file_t=0.234"
julia> eval(Meta.parse(strg))
0.234
julia> file_t
0.234

```

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2020, 1:15pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/6 "2020-07-15T13:15:51Z")

</div>

[https://github.com/JuliaIO/Formatting.jl/issues/21](https://github.com/JuliaIO/Formatting.jl/issues/21)

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2020, 1:16pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/7 "2020-07-15T13:16:12Z")

</div>

[https://gist.github.com/c42f/9999dc6f9b63a9bd4ea4237a95876475](https://gist.github.com/c42f/9999dc6f9b63a9bd4ea4237a95876475)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 15, 2020, 1:24pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/8 "2020-07-15T13:24:28Z")

</div>

> [@erlebach](#):
>
> What solutions does Julia offer besides regular expressions?

Any particular reason you want to avoid regexps? They are often a good solution for tasks like this.

To compare with @oheil’s nice solution,

```julia
using BenchmarkTools
strg="time t=0.234_ext.bson"
f1(str) = parse(Float64, split(split(str, '=')[2], '_')[1])
f2(str) = parse(Float64, match(r".*=(.*)_.*", str).captures[1])

julia> @btime f1($strg)
  592.303 ns (6 allocations: 448 bytes)
0.234

julia> @btime f2($strg)
  389.634 ns (4 allocations: 288 bytes)
0.234

```

In addition, it is easier to do more validation wiht regexps.

---

<div class="post-metadata">

**Author:** ![erlebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/erlebach/32/12973_2.png) [@erlebach](https://discourse.julialang.org/u/erlebach)\
**Post date:** [July 15, 2020, 1:52pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/9 "2020-07-15T13:52:25Z")

</div>

I have more than enough information now, thanks.  
I tend to avoid regular expressions because I never learned them really well. I have used them in several languages but always forget the subtleties. I recognize their value.

But I also believe if efficiency is not an issue (reading data once for example, and small amounts of data), I feel that regex is overkill and simpler solutions should be available.  
Of course, nothing prevents me from writing specialized routines that are easy to use and tuned to my workflow. It all takes time.

I appreciate the input and the help to make me more proficient.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [July 15, 2020, 2:18pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/10 "2020-07-15T14:18:01Z")

</div>

Perhaps a bit simpler would be to just look for the first float-like portion in the string:

```julia
f3(str) = parse(Float64, match(r"\d*\.?\d+", str).match)

```

Happens to be a bit faster, but more importantly (IMO) I think it’s easier to read and should be more robust.

---

<div class="post-metadata">

**Author:** ![ellocco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ellocco/32/31331_2.png) [@ellocco](https://discourse.julialang.org/u/ellocco)\
**Post date:** [October 13, 2022, 8:25am UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/11 "2022-10-13T08:25:33Z")

</div>

if your string may contain exponential numbers you might modify the inner part as follows:

```julia
match(r"\d*\.?\d+(e[+|-]\d*)?", str).match

```

(I hope this is correct, I have not tested all possible cases)

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 13, 2022, 9:14am UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/12 "2022-10-13T09:14:13Z")

</div>

> [@ellocco](#):
>
> `match(r"\d*\.?\d+(e[+|-]\d*)?", str).match`

It doesn’t accurately match this:

```julia
julia> match(r"\d*\.?\d+(e[+|-]\d*)?", "2.3e6").match
"2.3"

```

Nor `Float32`:

```julia
julia> match(r"\d*\.?\d+(e[+|-]\d*)?", "2.3f-6").match
"2.3"

```

Changing `"e[+|-]"` to `"[ef][+-]?"` helps:

```julia
julia> match(r"\d*\.?\d+([ef][+-]?\d*)?", "2.3e6").match
"2.3e6"

julia> match(r"\d*\.?\d+([ef][+-]?\d*)?", "2.3f-6").match
"2.3f-6"

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 13, 2022, 9:40am UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/13 "2022-10-13T09:40:23Z")

</div>

It needs `-?`, in order to catch negative floats as well:

```julia
match(r"-?\d*\.?\d+([ef][+-]?\d*)?", "-2.3f-6").match

```

---

<div class="post-metadata">

**Author:** ![ellocco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ellocco/32/31331_2.png) [@ellocco](https://discourse.julialang.org/u/ellocco)\
**Post date:** [October 13, 2022, 1:00pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/14 "2022-10-13T13:00:43Z")

</div>

> [@DNF](#):
>
> ```julia
> julia> match(r"\d*\.?\d+([ef][+-]?\d*)?", "2.3f-6").match
> "2.3f-6"
> 
> ```

comment:  
parse cannot handle “f+/-xx”

```julia
parse(Float64, "-1.234f-4")
ERROR: ArgumentError: cannot parse "-1.234f-04" as Float64

```

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 13, 2022, 1:02pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/15 "2022-10-13T13:02:19Z")

</div>

That’s probably because it’s not a `Float64`, but a `Float32`.

---

<div class="post-metadata">

**Author:** ![ellocco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ellocco/32/31331_2.png) [@ellocco](https://discourse.julialang.org/u/ellocco)\
**Post date:** [October 13, 2022, 1:03pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/16 "2022-10-13T13:03:57Z")

</div>

```julia
parse(Float32, "-1.234f-04")

```

fails as well ☹

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 13, 2022, 1:06pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/18 "2022-10-13T13:06:33Z")

</div>

> [@erlebach](#):
>
> I tend to avoid regular expressions because I never learned them really well.

See [https://regexr.com](https://regexr.com) (I believe the site I had in mind), but also consider: [GitHub - jkrumbiegel/ReadableRegex.jl: regexes for people who don't really want to learn or read regexes](https://github.com/jkrumbiegel/ReadableRegex.jl)

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 13, 2022, 1:12pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/19 "2022-10-13T13:12:39Z")

</div>

I’m stumped. A parser bug?

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 13, 2022, 1:15pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/20 "2022-10-13T13:15:51Z")

</div>

See [this issue here](https://github.com/JuliaLang/julia/issues/5690), confirming that we cannot parse as Float32 strings like `"-1.234f04"` using ` parse(Float32, "-1.234f04")`

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 13, 2022, 1:26pm UTC](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126/21 "2022-10-13T13:26:05Z")

</div>

`Meta.parse("-1.234f04")` works.

[Next page](https://discourse.julialang.org/t/extracting-a-float-from-a-string/43126.md?page=2)
