I’ve been trying to add speed to XLSX.jl as it reads data from Excel in response to this issue.
I’ve been working with Claude and achieving some good results but one suggestion it made surprised me. It recommended replacing parse with what it calls “Clinger’s exact fast path” for a subset of possible Floats. Specifically:
For a plain decimal (-?digits[.digits][e±digits]) with at most 15 significant digits and a decimal exponent p with |p| ≤ 22:
- the digits form an integer m < 10¹⁵ < 2⁵³, which is exact in Float64;
- 10^|p| is also exact, since 5²² < 2⁵³;
- so the single IEEE operation m * 10^p (or m / 10^-p) is correctly rounded, which is the same result a correct
strtod(the C library method Claude says Base uses) returns.
It puts everything else through the standard path in Base unchanged.
In Windows 11, it measures this as follows:
| Input | parse(Float64, s) | fast path |
|---|---|---|
| typical values, e.g. “1234.5678” | 323 ns | 13.6 ns (24×) |
| 17-significant-digit values (fall through) | 403 ns | 396 ns (no overhead) |
I have always understood that Base was pretty well optimised, so my naïve reasoning is that, if Claude is right, Base must be handling a range of possible inputs that do not arise with Excel, so can be ignored in XLSX.jl. I have no idea if this is right.
So, three questions:
- Is Claude’s proposal credible?
- Will similar improvements be seen on Linux/MacOS and other supported systems?
- Should I consider including this change?
Any advice very welcome because this is well outside my expertise!