# Question about parse(Float) in Base

**URL:** <https://discourse.julialang.org/t/question-about-parse-float-in-base/139812>\
**Category:** Performance\
**Tags:** question, parser, xlsx\
**Created:** [October 4, 2026, 6:52am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812 "2026-10-04T06:52:19Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [October 4, 2026, 6:52am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/1 "2026-10-04T06:52:20Z")

</div>

I’ve been trying to add speed to XLSX.jl as it reads data from Excel in response to [this issue](https://github.com/JuliaData/XLSX.jl/issues/462).

I’ve been working with Claude and achieving some good results but one suggestion it made surprised me. It recommended replacing `parse` with what it calls “Clinger’s exact fast path” for a subset of possible Floats. Specifically:

> For a plain decimal (-?digits[.digits][e±digits]) with at most 15 significant digits and a decimal exponent p with |p| ≤ 22:
> 
> - the digits form an integer m \< 10¹⁵ \< 2⁵³, which is exact in Float64;
> - 10^|p| is also exact, since 5²² \< 2⁵³;
> - so the single IEEE operation m \* 10^p (or m / 10^-p) is correctly rounded, which is the same result a correct `strtod` (the C library method Claude says Base uses) returns.

It puts everything else through the standard path in Base unchanged.

In Windows 11, it measures this as follows:

| Input | parse(Float64, s) | fast path |
| --- | --- | --- |
| typical values, e.g. “1234.5678” | 323 ns | 13.6 ns (24×) |
| 17-significant-digit values (fall through) | 403 ns | 396 ns (no overhead) |

I have always understood that Base was pretty well optimised, so my naïve reasoning is that, if Claude is right, Base must be handling a range of possible inputs that do not arise with Excel, so can be ignored in XLSX.jl. I have no idea if this is right.

So, three questions:

- Is Claude’s proposal credible?
- Will similar improvements be seen on Linux/MacOS and other supported systems?
- Should I consider including this change?

Any advice very welcome because this is well outside my expertise!

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [October 4, 2026, 7:26am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/2 "2026-10-04T07:26:25Z")

</div>

All I can say is that the argument seems a bit incoherent. If you consider all the significant digits as an integer, you need to shift the exponent correspondingly to compensate for decimal positions. But the shifted exponent wouldn’t stay within ±22, which the rest of the analysis assumes.

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [October 4, 2026, 7:50am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/3 "2026-10-04T07:50:56Z")

</div>

It’s possible I’ve over simplified too far. Claude gave me an extensive explanation but I didn’t paste it here verbatim because to do so would probably violate forum rules. It is hard to summarise something you don’t fully understand!

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [October 4, 2026, 8:07am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/4 "2026-10-04T08:07:09Z")

</div>

In hopes of escaping censure:

**Title:** parse(Float64, s) is ~24x slower than an exact fast path on Windows. Worth adding to Base?

While fixing a read-performance regression in XLSX.jl (JuliaData/XLSX.jl#462), I found that parse(Float64, s) was the largest single cost of reading a numeric spreadsheet: 16–19% of readtable time in a profile. In 1.13, Base’s method is a ccall to jl\_try\_substrtod, the C library’s strtod.

I added Clinger’s exact fast path in front of it. For a plain decimal (-?digits[.digits][e±digits]) with at most 15 significant digits and a decimal exponent p with |p| ≤ 22:

- the digits form an integer m \< 10¹⁵ \< 2⁵³, which is exact in Float64;
- 10^|p| is also exact, since 5²² \< 2⁵³;
- so the single IEEE operation m \* 10^p (or m / 10^-p) is correctly rounded, which is the same result a correct strtod returns.

Anything else falls through to parse(Float64, s) unchanged: more digits, larger exponents, whitespace, inf/nan, hex, or malformed input. So results and errors outside the fast path are identical. The code is a ~60-line loop over the code units, with no allocation: [link to \_fast\_decimal in src/cell.jl on the branch]

**Measurements** (Windows 11, x86\_64, Julia 1.13.1, SubString{String} inputs):

| **Input** | **parse(Float64, s)** | **fast path** |
| --- | --- | --- |
| typical values, e.g. “1234.5678” | 323 ns | 13.6 ns (24x) |
| 17-significant-digit values (fall through) | 403 ns | 396 ns (no overhead) |

**Correctness.** Over 3 million random strings, 2.06 million took the fast path, with 0 mismatches against parse (compared with reinterpret(UInt64, …)). XLSX.jl’s tests add 1.5 million strings in the forms Excel writes, every numeric cell in its test files, and edge cases (-0, 1E22 vs 1E23, 15 vs 16 digits, .5, 1., rejected inputs).

**End to end.** With only this change, XLSX.jl’s readtable on a 30,000 × 84 Excel-style sheet went from 3.14 s to 2.60 s (−17%), and from 4.61 s to 3.68 s (−20%) on a 50,000 × 70 sheet.

**Questions:**

1. Is strtod known to be slow on Windows in particular? I’ve only measured there. Linux and macOS numbers would help show whether this is general.
2. Would Base take a pure-Julia fast path like this ahead of the ccall? Or go further, with something like Eisel–Lemire (as in fast\_float, Rust and, I believe, Parsers.jl), which covers 16–19 digits as well? Many real values are 16–17-digit computed results.
3. Is there a reason for keeping C strtod that I’m missing (locale, error semantics, platform behaviour)?

**The fast path relies on IEEE binary64 arithmetic without x87 extended precision. My understanding is that this holds on all of Julia’s supported platforms (SSE2 on i686).**

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 4, 2026, 8:39am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/5 "2026-10-04T08:39:39Z")

</div>

Parsers.jl provides optimized numeric parsers that work directly on `Vector{UInt8}` buffers and views thereof without going through strings.

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [October 4, 2026, 8:48am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/6 "2026-10-04T08:48:18Z")

</div>

But strings are what Excel provides - they’re the starting point. (Edit: Maybe I’ve misunderstood. Need to check this out with XML.jl)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 4, 2026, 8:51am UTC](https://discourse.julialang.org/t/question-about-parse-float-in-base/139812/7 "2026-10-04T08:51:06Z")

</div>

You can also use Parsers.jl on a string, or the low-level functions on `codeunits(string)`, of course. But you’re probably losing performance by allocating a bunch of individual `String` objects just to parse them.
