# Why does the speed of Julia in AOT compilation differ from UX4?

**URL:** <https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657>\
**Category:** Performance\
**Tags:** performance, compilation, benchmark\
**Created:** [September 20, 2024, 8:12pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657 "2024-09-20T20:12:53Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![aaoo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaoo/32/209309_2.png) [@aaoo](https://discourse.julialang.org/u/aaoo)\
**Post date:** [September 20, 2024, 8:12pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/1 "2024-09-20T20:12:53Z")

</div>

As the following picture shows, there are two different Julia versions in the speed comparison. The AOT compiled version reaches speeds as fast as Fortran, etc. However, the other version is slower and closer to Python.

I have a couple of questions:

1. Could anyone explain what ux4 mean?
2. For ordinary users, if they follow the guidance of [Performance Tips](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-tips), which version of speed are they most likely to achieve? And how fast?

 ![combined_results](https://global.discourse-cdn.com/julialang/original/3X/b/f/bf7600d8fc0ba57f3c3790b5b4352cdd994db828.png)

picture source: [Latest - Speed comparison](https://niklas-heer.github.io/speed-comparison/)

---

<div class="post-metadata">

**Author:** ![nsajko](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nsajko/32/221187_2.png) [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Post date:** [September 20, 2024, 8:30pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/2 "2024-09-20T20:30:36Z")

</div>

This is a question you should pose to whoever made that “speed comparison”. In general Julia can be as fast as C++ or Rust, and faster than C. Keep in mind good performance may require effort, in any language.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 20, 2024, 8:36pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/3 "2024-09-20T20:36:35Z")

</div>

Those two data points seem to be comparing two different implementations, one in [`leibniz.jl`](https://github.com/niklas-heer/speed-comparison/blob/4a7721f6d432ddf4d6ba43291d10b1f4aeca766b/src/leibniz.jl) and one in [`leibniz_ux4.jl`](https://github.com/niklas-heer/speed-comparison/blob/master/src/leibniz_ux4.jlt). The latter unrolls the inner loop 4x (which reduces overhead).

However, it it looks like they are making the usual mistake of including the startup and compilation time for Julia in the benchmark. i.e. [just timing `julia script.jl` from the shell](https://github.com/niklas-heer/speed-comparison/blob/4a7721f6d432ddf4d6ba43291d10b1f4aeca766b/Earthfile#L257).

So, basically the results are misleading for how the performance would be in a more realistic program that does heavy number crunching (since you only care about the performance if it takes much longer than a second to run, and in this case the one-time compilation overhead is negligible). It’s also a bit unfair to compare it to AOT compiled languages like C, since you’re not including the time to run the compiler for those language.

This also makes the benefit of loop unrolling minimal, since they are mostly measuring startup time.

This comes up every single time people do a cross-language benchmark, and it’s a bit tiring to address over-and-over again.

> <https://github.com/niklas-heer/speed-comparison/issues/130>
>
> In Julia, and perhaps in some of the other languages, you are mostly measuring t…he one-time startup cost, rather than the cost of actually doing the mathematical calculation. So the results are misleading as a guide for speed of realistic number-crunching tasks, since compute-heavy problems inevitably run for long enough that the startup time is negligible.
> 
> In Julia's case, this startup cost is especially heavy because it includes a one-time cost of compiling the function you are benchmarking. It would be like including the time for the C compiler in the C benchmark.
> 
> Recommendation: do the timing \*within\* each language's script (all languages offer high-resolution timers), and print out the elapsed compute time for calling \`f(rounds)\` \*only\*. In Julia's case, call the function once (e.g. with \`rounds=1\`) before beginning the timing, to trigger its compilation.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 20, 2024, 8:43pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/4 "2024-09-20T20:43:51Z")

</div>

This is how the two benchmarks are executed.

```julia
julia:
  # We have to use a special image since there is no Julia package on alpine 🤷‍♂️
  FROM julia:1.8.2-alpine3.16
  DO +PREPARE_ALPINE
  DO +ADD_FILES --src="leibniz.jl"
  DO +BENCH --name="julia" --lang="Julia" --version="julia --version" --cmd="julia leibniz.jl"

julia-compiled:
  # We need the Debian version otherwise the build doesn't work
  FROM julia:1.8.2
  DO +PREPARE_DEBIAN
  RUN apt-get update && apt-get install -y gcc g++ build-essential cmake
  DO +ADD_FILES --src="leibniz_compiled.jl"
  COPY ./src/leibniz.jl ./
  RUN julia -e 'using Pkg; Pkg.add(["StaticCompiler", "StaticTools"]); using StaticCompiler, StaticTools; include("./leibniz_compiled.jl"); compile_executable(mainjl, (), "./")'
  DO +BENCH --name="julia-compiled" --lang="Julia (AOT compiled)" --version="julia --version" --cmd="./mainjl"

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 20, 2024, 8:44pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/5 "2024-09-20T20:44:44Z")

</div>

Oh, I see that there is a “Julia (AOT Compiled)” line, which removes some of the startup cost.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 20, 2024, 8:46pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/6 "2024-09-20T20:46:05Z")

</div>

Yes, but it uses `StaticCompiler.jl` which is an unrealistic scenario for most users.

Most realistic would be either a pre-compiled package, or a program compiled with PackageCompiler.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 20, 2024, 8:51pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/7 "2024-09-20T20:51:16Z")

</div>

And in any case, why include misleading results (startup-dominated timings) at all in the benchmark? I filed an issue suggesting that they should really just measure the compute time _within_ each language’s code.

But this shows up over and over; it seems kind of pointless to try to correct every random amateur benchmarking attempt.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [September 20, 2024, 11:02pm UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/8 "2024-09-20T23:02:54Z")

</div>

> [@stevengj](#):
>
> ; it seems kind of pointless to try to correct every random amateur benchmarking attempt.

💯 this

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [September 21, 2024, 12:20am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/9 "2024-09-21T00:20:01Z")

</div>

Relevant JuliaCon talk about this benchmark:

[![](https://global.discourse-cdn.com/julialang/original/3X/7/b/7b16251c6c102497c3e5b08a2b89353a961c959c.jpeg "The Slack thread that would not die | Miguel Raz Guzmán Macedo | JuliaCon 2023") ](https://www.youtube.com/watch?v=iEaXqu1niXQ)

---

<div class="post-metadata">

**Author:** ![aaoo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaoo/32/209309_2.png) [@aaoo](https://discourse.julialang.org/u/aaoo)\
**Post date:** [September 21, 2024, 1:46am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/10 "2024-09-21T01:46:59Z")

</div>

The AOT compiled version is as fast as Fortran. A realistic question is: for ordinary users, how could the speed of the code that not utilize `StaticCompiler. jl` close to the AOT version?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 21, 2024, 2:00am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/11 "2024-09-21T02:00:17Z")

</div>

The benchmark is not measuring the speed of the numerical code in the non-AOT case, it is mostly measuring startup time. And that is totally an artificial consequence of the fact that the benchmark is so short (a fraction of a second), so that startup time dominates.

The actual numerical calculation will be the same speed in the AOT and non-AOT cases.

For ordinary users, if you care about numerical performance, it’s probably for code that takes more than a few seconds to run. In which case startup time is irrelevant

---

<div class="post-metadata">

**Author:** ![aaoo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaoo/32/209309_2.png) [@aaoo](https://discourse.julialang.org/u/aaoo)\
**Post date:** [September 21, 2024, 3:09am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/12 "2024-09-21T03:09:14Z")

</div>

Thank you and everyone else for the reply. 🙏

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [September 21, 2024, 3:14am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657/13 "2024-09-21T03:14:26Z")

</div>

> [@aaoo](#):
>
> for ordinary users

Depends on what the user is doing. Are they repeatedly running a script from the command line, thus reloading and recompiling the sysimage and the script’s code and imports? Then the startup time is going to add up. Does the script have a parameter for more iterations or are they working within a session? Then the startup time only needs to happen once. When people argue about what’s fair to include in a benchmark, they’re often arguing over workflow like this. The AOT benchmark shows that when the startup and overhead for Julia’s interactivity is omitted, the particular algorithm runs as fast as other compiled languages.

Ordinary _Julia_ users are repeating calls in the Julia session, not using StaticCompiler to make an minimal executable. StaticCompiler trims so much overhead that many Julia features can’t be supported. juliac (development for v1.12 recently announced) and SyslabCC (proprietary, usable) aim to support more of Julia, but there is a fundamental feature-overhead tradeoff.
