Julia faster than C++ on a converted large simulation model

So, this morning I have asked Claude to convert our 12 years of development, large C++ simulation model in julia.
I have still to check it properly of course, but these are the initial outcomes. Very surpriced to see Julia much faster/efficient than C++:slight_smile:

What if you asked Claude to improve the performance of the C++ code as well? :wink:

Then repeat.

I calculate about 6.3x speedup, great, but not largest I’ve seen for AI rewrite! Even for $10.64 with EvoX Genesis. There seemingly 6.41× speedup is correct, but unfair comparison to your speedup, since only a subcomponent (you might have even higher speeup somewhere, have you dug into where and why faster?), and the “End-to-end benchmark: .. runs 1.25–1.32× faster than Fortran with exact integrator parity” seems more justifiably comparable to your speedup.

Can you say which Claude? Like Claude Code presumably, and Opus 5.5 and/or Fable 5.1? I’m not sure you can use Claude code with two models at the same time as with EvoX (or Devin Fusion). Do you know cost or tokens spent?

And do you want to say anything about code size? What does that 3.0 GB refer to, max RAM use?

1.55–6.87× median speedup, Fortran → Rust

89,946 lines of Rust replacing 139,414 lines of Fortran; 1,052 tests with 0 failures; a median speedup of 1.55–6.87×.

At GitHub I see:

Highlights

  • ~67K LOC of pure Rust library code (129 files, 13 crates), #![forbid(unsafe_code)] everywhere — zero unsafe, zero FFI.
  • ~1050 tests green (cargo test --workspace → 1052 passed / 0 failed / 18 ignored), including golden tests ported from the original Fortran test_output files (bit-exact or ~1e-9).
  • No LAPACK/BLAS FFI — matrix solvers (dense/banded/tridiag/block-tridiag/ sparse LU), ODE integrators (dopri5, dop853, ros2, rodas*), bicubic/bipm interpolation, and the approx21 network with analytic Jacobian are all pure-Rust ports.
  • End-to-end benchmark: the net 1-zone burn pipeline (approx21, the project’s end-to-end proxy) runs 1.25–1.32× faster than Fortran with exact integrator parity (nstep 199 / naccpt 198 / nrejct 1 / nfcn 8889 / njac 199).

“Every workload is faster than the original Fortran.” though seemingly fastest only | num_newton (10K solves) | 0.057918 | 0.009029 | 6.41× | bit-exact | and all other workloads are measured in seconds, not all bit-exact, but with acceptable tolerance.

Maybe you can ask Claude the main reason why the new version is faster.

Why the Julia version is so faster than the C++ version?

It isn’t really Julia being faster than C++ as a language. The speed comes from how the data is accessed, which the C++ code does very expensively. The port went faster in two steps:

Version default (100 years)
C++ 32m40s
First Julia version: line-by-line translation 16m33s
Final Julia version: after profiling and caching 5m12s

1. Data look-ups. Every gpd/gfd call in C++ builds a key string like "pl#11001#hardWRoundW#" and searches a std::map, a tree compared string by string. Most product queries are also prefix scans. These calls sit inside the pixel × forest type × diameter class × product loops, so there are millions per year.

  • First Julia version: it already used hash tables for exact keys and cached the prefix-scan results.
  • Final version: it caches by key components, so no string is built at all. It also computes regional values once per region rather than once per pixel; they can’t change inside those loops.

It has to be said that using hash tables in Julia is much easier than in C++, at least at the time I started coding it, so this is another dimension that make it easier to “optimize” in julia…