# Julia vs C++ speed

**URL:** <https://discourse.julialang.org/t/julia-vs-c-speed/67219>\
**Category:** General Usage\
**Tags:** performance, c\
**Created:** [August 28, 2021, 2:24pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219 "2021-08-28T14:24:14Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 2:24pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/1 "2021-08-28T14:24:14Z")

</div>

Hi!

I am currently doing a very “nice” project of converting algorithms written in Julia to C++ 😢 I need those algorithms in an embedded environment and, unfortunately, Julia cannot generate compiled code that can be embedded yet.

I am taking a very nice approach with the help of CxxWrap.jl. The translation is happening smoothly since I can rewrite function by function and test if anything breaks.

All the algorithms are the ones in [SatelliteToolbox.jl](https://github.com/JuliaSpace/SatelliteToolbox.jl).

The **very** interesting part was when I started to compare the performance. There are three algorithms that are relatively costly for the satellite computer: IGRF (geomagnetic field), SGP4 (orbit propagator), and FK5 reduction (nutation calculation). Check out the comparison between the Julia version in SatelliteToolbox.jl and the C++ version (using `-O3 -march=native`):

| Algorithm | Julia [ns] | C++ [ns] |
| --- | --- | --- |
| IGRF | 1373 | 1425 |
| SGP4 propagation | 340 | 660 |
| FK5 nutation | 1436 | 1201 |

Notice that Julia “lost” only in the nutation computation, which is an algorithm that has to go through a table of 106 lines and 9 columns and perform sums (I think there is room for optimization here). I am also not sure if anything will help if I recompile Julia with `-O3 -march=native`.

My colleagues and I were just astonished with those results. This is a language that seems interpreted vs C++. I always thought that pursuing C++ speed was the ultimate goal, but now I see that Julia can be even faster for reasons I just cannot explain 😃

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 2:31pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/2 "2021-08-28T14:31:07Z")

</div>

> [@Ronis\_BR](#):
>
> nutation

> <https://github.com/JuliaSpace/SatelliteToolbox.jl/blob/abd462a68a498f4ee5cd2b8f9c96e50a2b59d203/src/transformations/fk5/nutation.jl#L280>

don’t do this, this copies every single loop. Better just don’t make the look up table in Matrix to begin with, just have these columns separately in `SVector`? That way you also avoid promoting integers to floats in the Matrix.

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 2:42pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/3 "2021-08-28T14:42:15Z")

</div>

> [@jling](#):
>
> don’t do this, this copies every single loop. Better just don’t make the look up table in Matrix to begin with, just have these columns separately in `SVector` ? That way you also avoid promoting integers to floats in the Matrix.

I really need help here. In my local branch, I switched the dimensions since Julia is column-major. The benchmark were obtained in this configuration. Also tried to make `nut_coefs_1980` a vector of vectors. The speed was worse. I also converted it to a vector of `SVector`. In this case, the performance was almost the same as when I used the look-up matrix but with dimensions permuted.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 2:43pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/4 "2021-08-28T14:43:10Z")

</div>

any runnable snippet with that function alone? I can give it a stab

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 2:43pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/5 "2021-08-28T14:43:36Z")

</div>

Sure:

```julia
using SatelliteToolbox
nutation_fk5(2.460115315972222e6)

```

Thanks!!

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [August 28, 2021, 2:45pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/6 "2021-08-28T14:45:18Z")

</div>

> [@jling](#):
>
> don’t do this, this copies every single loop.

[`nut_coefs_1980` is a 2D matrix of Floats](https://github.com/JuliaSpace/SatelliteToolbox.jl/blob/abd462a68a498f4ee5cd2b8f9c96e50a2b59d203/src/transformations/fk5/nutation.jl#L48-L54), there’s no allocation happening here. Just regular array access/load from memory. Since the table in source code mixes integers and floats, the integers will be promoted to floats as well, making sure the matrix is uniformly typed.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 2:46pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/7 "2021-08-28T14:46:31Z")

</div>

oops, I thought it was slicing, now I realized it’s accessing a single element at a time.

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 2:47pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/8 "2021-08-28T14:47:30Z")

</div>

Ah @jling notice that this is not my newest code. I did not pushed because I am doing some tests. The changes were:

```julia
const nut_coefs_1980 = permutedims([
...
])

```

and

```julia
        an1 = nut_coefs_1980[1,i]
        an2 = nut_coefs_1980[2,i]
        an3 = nut_coefs_1980[3,i]
        an4 = nut_coefs_1980[4,i]
        an5 = nut_coefs_1980[5,i]
        Ai = nut_coefs_1980[6,i]
        Bi = nut_coefs_1980[7,i]
        Ci = nut_coefs_1980[8,i]
        Di = nut_coefs_1980[9,i]

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 2:48pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/9 "2021-08-28T14:48:28Z")

</div>

btw,

```julia
@turbo for i = 1:n_max

```

seem to work:

```julia
julia> @benchmark nutation_fk5(2.460115315972222e6)
BenchmarkTools.Trial: 10000 samples with 10 evaluations.
 Range (min … max): 1.166 μs … 3.456 μs ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 1.183 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 1.188 μs ± 45.902 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

      █▁                                                      
  ▂▄████▄▃▂▂▂▂▂▂▂▂▁▁▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▂
  1.17 μs Histogram: frequency by time 1.4 μs <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

after:

```julia
julia> @benchmark nutation_fk5(2.460115315972222e6)
BenchmarkTools.Trial: 10000 samples with 219 evaluations.
 Range (min … max): 339.320 ns … 436.995 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 348.790 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 349.821 ns ± 5.359 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

             ▂▅▇▇█▆▄▁                                            
  ▂▂▂▂▂▂▃▃▄▅▇████████▇▅▄▃▃▃▃▂▂▂▂▂▂▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂ ▃
  339 ns Histogram: frequency by time 377 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

results are the same within float point error I think:

```julia
julia> nutation_fk5(2.460115315972222e6)
#n_max = 106
(0.4090395463527846, 3.477881920473464e-5, -4.185255450109059e-5)

julia> nutation_fk5(2.460115315972222e6)
(0.4090395463527846, 3.477881920473462e-5, -4.185255450109061e-5)

```

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 2:55pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/10 "2021-08-28T14:55:58Z")

</div>

This is AWESOME! Thanks! I did not know about `@turbo` before!

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 3:04pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/11 "2021-08-28T15:04:04Z")

</div>

With @jling 's suggestion the table can be updated:

| Algorithm | Julia [ns] | C++ [ns] |
| --- | --- | --- |
| IGRF | 1373 | 1425 |
| SGP4 propagation | 340 | 660 |
| FK5 nutation | **409** | 1201 |

The only “downside” is that the compilation time is much higher when using `@turbo`. Also, I think we can perform some similar tuning into the C++ code as well.

Anyway, the fact that straightforward Julia code can match and even beat C++ is just amazing!

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 28, 2021, 3:08pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/12 "2021-08-28T15:08:32Z")

</div>

> [@Ronis\_BR](#):
>
> The only “downside” is that the compilation time is much higher when using `@turbo` .

test it with only

```julia
@inbounds @simd for i = 1:n_max

```

(may it was tested). It will compile faster, and sometimes the difference to `@turbo` is not that much.

edit: in this case it will probably make a difference, because you have a `sincos` inside the loop and, if I’m not mistaken, `@turbo` uses faster versions of these functions. You may want to try `@fastmath` on that function, if that is the case. (in both cases those faster versions have some accuracy loss, I don’t know if that is relevant for the satellites 🙂 .

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 3:14pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/13 "2021-08-28T15:14:09Z")

</div>

> [@Ronis\_BR](#):
>
> IGRF

now I wonder if similar practice can be done:

> <https://github.com/JuliaSpace/SatelliteToolbox.jl/blob/2adbf245bb710e2808939b4898ea936bfefb374f/src/earth/geomagnetic_field_models/igrf/igrf.jl#L445>

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [August 28, 2021, 3:14pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/14 "2021-08-28T15:14:41Z")

</div>

> [@lmiq](#):
>
> in both cases those faster versions have some accuracy loss, I don’t know if that is relevant for the satellites 🙂

Famous last words 😄

(Well, famous last words at a job…)

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 28, 2021, 3:17pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/15 "2021-08-28T15:17:40Z")

</div>

Thanks for the tips! I got those values:

1. `@simd`: 1409 ns (slightly faster).
2. `@fastmath`: 1386 ns (getting there).
3. `@fastmath @simd`: 1380 ns.

> [@lmiq](#):
>
> (in both cases those faster versions have some accuracy loss, I don’t know if that is relevant for the satellites

We cannot loose accuracy here. However, all those versions have exactly the same result.

> [@lmiq](#):
>
> because you have a `sincos` inside the loop and, if I’m not mistaken, `@turbo` uses faster versions of these functions.

If there is a faster `sincos` version, why Julia does not use it?

> [@jling](#):
>
> now I wonder if similar practice can be done:

Probably! The only problem with IGRF is that I compute sines and cosines recursively. Hence, we do not have that assumption in which we can change the loop order.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 28, 2021, 3:19pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/16 "2021-08-28T15:19:29Z")

</div>

> [@lmiq](#):
>
> in both cases those faster versions have some accuracy loss, I don’t know if that is relevant for the satellites

you say accuracy loss, but they are actually more accurate than loop `+`, just slightly less accurate than `sum`

```julia
julia> a = rand(Float16, 10000);

julia> sum(a)
Float16(5.024e3)

julia> foldl(+, a)
Float16(2.048e3)

julia> let res = zero(Float16)
           @turbo for i in 1:length(a)
               res += a[i]
           end
           res
       end
5025.9033f0

```

---

<div class="post-metadata">

**Author:** ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)\
**Post date:** [August 28, 2021, 3:25pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/17 "2021-08-28T15:25:23Z")

</div>

I think the `@turbo` sincos is slightly less precise. I don’t know how much is slightly less, probably check [https://github.com/JuliaSIMD/SLEEFPirates.jl](https://github.com/JuliaSIMD/SLEEFPirates.jl).

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 28, 2021, 3:29pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/18 "2021-08-28T15:29:01Z")

</div>

> [@Ronis\_BR](#):
>
> If there is a faster `sincos` version, why Julia does not use it?

Because of the precision. It is like using fastmath compiler options in general. For some applications the difference is important (not any of mine).

---

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [August 28, 2021, 3:40pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/19 "2021-08-28T15:40:02Z")

</div>

> [@Ronis\_BR](#):
>
> If there is a faster `sincos` version, why Julia does not use it?

It’s not necessarily faster in general (for scalar arguments). Instead, it’s optimized for SIMD vectorization, where it outperforms the `Base.sincos` version (which is optimized for scalars). For example,

```julia
julia> using LoopVectorization, BenchmarkTools

julia> @benchmark Base.sincos($(Ref(1.0))[])
BenchmarkTools.Trial: 10000 samples with 999 evaluations.
 Range (min … max): 9.604 ns … 35.111 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 9.608 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 9.699 ns ± 1.163 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▇█▁ ▁
  ███▅▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▇▇ █
  9.6 ns Histogram: log(frequency) by time 9.82 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark LoopVectorization.SLEEFPirates.sincos($(Ref(1.0))[])
BenchmarkTools.Trial: 10000 samples with 997 evaluations.
 Range (min … max): 18.468 ns … 69.659 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 18.827 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 18.912 ns ± 1.597 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

              █                                                
  ▂▅▃▂▁▁▁▁▁▁▂▄█▆▂▁▂▁▁▁▁▁▁▁▁▂▂▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂▂▂▂▂▂▁▁▁▁▂▂ ▂
  18.5 ns Histogram: frequency by time 20.2 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> function foo_simd!(si, co, x)
           @inbounds @simd ivdep for i in eachindex(x)
               si[i], co[i] = sincos(x[i])
           end
       end
foo_simd! (generic function with 1 method)

julia> function foo_turbo!(si, co, x)
           @turbo for i in eachindex(x)
               si[i], co[i] = sincos(x[i])
           end
       end
foo_turbo! (generic function with 1 method)

julia> x = randn(10^3); si = similar(x); co = similar(x);

julia> @benchmark foo_simd!($si, $co, $x)
BenchmarkTools.Trial: 10000 samples with 4 evaluations.
 Range (min … max): 7.246 μs … 23.578 μs ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 7.365 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 7.425 μs ± 613.371 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

   ▇█ ▁
  ▄██▄▇▄▁▅█▆▃▁▃▄▃▃▃▁▁▃▁▁▁▃▁▃▁▁▁▃▁▁▁▁▃▁▃▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▆ █
  7.25 μs Histogram: log(frequency) by time 10.6 μs <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark foo_turbo!($si, $co, $x)
BenchmarkTools.Trial: 10000 samples with 10 evaluations.
 Range (min … max): 1.981 μs … 6.003 μs ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 2.006 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 2.017 μs ± 151.188 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

         ▁█▇                                                   
  ▂▃▃▃▂▁▂███▃▂▂▂▁▂▂▂▁▁▁▁▂▂▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂▁▁▁▁▁▁▁▁▁▁▁▂▁▁▁▂▂▂ ▂
  1.98 μs Histogram: frequency by time 2.15 μs <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> foo_simd!(si, co, x)

julia> si_new = similar(si); co_new = similar(co);

julia> foo_turbo!(si_new, co_new, x)

julia> si ≈ si_new
true

julia> co ≈ co_new
true

```

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [August 28, 2021, 4:23pm UTC](https://discourse.julialang.org/t/julia-vs-c-speed/67219/21 "2021-08-28T16:23:11Z")

</div>

Note `@turbo` will use `SLEEFPirates.sincos_fast` by default, but you could specify `SLEEFPirates.sincos`.

On speed/implementations differences- note that SIMD (Single Instruction Multiple Data) has to apply the same instruction to multiple data.  
This basically means that if your code has a branch, you have to take both sides of the branch and fuse the results at the end.

Sometimes, functions like `sincos` in scalar mode can be sped up by having efficient strategies for different segments, and then branching based on which segment you’re in. But you don’t want to do this for a SIMD version, you’d be best off with just 1 primary path through the code.

[Next page](https://discourse.julialang.org/t/julia-vs-c-speed/67219.md?page=2)
