# When Julia gets within 1-3x of C/C++ speed, why is C/C++ usually faster?

**URL:** <https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341>\
**Category:** Performance\
**Created:** [October 30, 2020, 10:51pm UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341 "2020-10-30T22:51:25Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [October 30, 2020, 10:51pm UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/1 "2020-10-30T22:51:26Z")

</div>

A notion that I’m getting from blogs and published papers comparing Julia with other languages in practical use cases is that Julia gets within 1-3x of the speed of the fastest implementation in C/C++ (though it’s worth mentioning that Julia is invariably more readable and does rarely beat C/C++). To be clear, I’m talking specifically about writers who made the effort to read the performance tips and do the things the Julia way: type stability, concrete fields, limited allocations. If they didn’t, they’d get a slowdown of several orders of magnitude, as many other topics here can demonstrate.

It seems that we can already give the compiler the sort of information that people broadly say is what makes compiled languages like C/C++ more efficient. The only big difference I can think of is Julia’s garbage collector, but I can’t say how that factors into anything because I don’t know any low-level languages. I’m hoping someone who does can give general reasons for the remaining bit of difference in performance.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 30, 2020, 11:08pm UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/2 "2020-10-30T23:08:05Z")

</div>

It’s very hard to speak about this in generalities because every individual little problem is actually a neverending rabbit hole of potential microoptimizations, special cases, and weird details you would never have guessed.

However, in general I would say that in my experience (as someone who reads a fair amount about this, but doesn’t use or know C) the two main reasons are

1. It’s awkward. You can basically write C flavoured julia code if you really want to, it’s just ugly and awkward. It feels like cheating if the goal is to compare julia to C, but then to write horrific, unidiomatic unsafe julia code that is basically just C.

2. Missing optimizations: There are just some optimizations that are possible in julia but not yet implemented because it’s hard or nobody has gotten around to it yet. Many of these things are missing optimizations in LLVM, or things that are awkward for us to communicate properly to LLVM. A great example of this would be vectorization and SIMD. It turns out you can squeeze some pretty insane performance out of julia loops if you use [https://github.com/chriselrod/LoopVectorization.jl](https://github.com/chriselrod/LoopVectorization.jl) and you can often completely smoke all but the most clever, handwritten custom assembly iteration schemes. This basically happens by bypassing LLVM’s looping stuff and getting Chris Elrod to do your code generation instead.

---

<div class="post-metadata">

**Author:** ![yuyichao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuyichao/32/20_2.png) [@yuyichao](https://discourse.julialang.org/u/yuyichao)\
**Post date:** [October 31, 2020, 12:04am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/3 "2020-10-31T00:04:40Z")

</div>

> [@Benny](#):
>
> The only big difference I can think of is Julia’s garbage collector,

If it’s within a factor of 3 than GC is no an issue.

---

<div class="post-metadata">

**Author:** ![daniel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/daniel/32/6480_2.png) [@daniel](https://discourse.julialang.org/u/daniel)\
**Post date:** [October 31, 2020, 1:50am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/4 "2020-10-31T01:50:18Z")

</div>

the wording of that last sentence made me chuckle

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 31, 2020, 1:56am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/5 "2020-10-31T01:56:55Z")

</div>

Only the finest handmade artisanal SIMD code from @Elrod.

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [October 31, 2020, 2:01am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/6 "2020-10-31T02:01:17Z")

</div>

With substantial software, when Julia gets within 1-3x of C/C++ speed, why has Julia obtained working, reliable, collaboratively written code 2-5x faster?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 31, 2020, 2:18am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/7 "2020-10-31T02:18:57Z")

</div>

I think the biggest reason for this is that to get the power of multiple dispatch, you would need to write all your code using C++ templates. Doing so to the extent Julia does would absolutely kill your compile times due to the lack of a JIT. Julia’s macros are also a huge part of the story here. Tools like LoopVectorization mean that idiots like me can write the equivalent to hand optimized assembly for any loop that is even a vague hotspot.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 31, 2020, 3:07am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/8 "2020-10-31T03:07:37Z")

</div>

Are there any plans to integrate LoopVectorization into Julia?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 31, 2020, 3:11am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/9 "2020-10-31T03:11:44Z")

</div>

Not in the short term. LoopVectorization is improving very quickly, and incorporating it into base would almost certainly massively slow progress. Furthermore, there is not a ton of benefit of adding it to Base.

---

<div class="post-metadata">

**Author:** ![KZiemian](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kziemian/32/9020_2.png) [@KZiemian](https://discourse.julialang.org/u/KZiemian)\
**Post date:** [November 2, 2020, 1:27am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/10 "2020-11-02T01:27:06Z")

</div>

I don’t know if this will help you, but Jeff Bazanson talks about Julica vs C, C++ speed around 16:45 of [_State of Julia_ Jeff Bezanson & Stefan Karpinski](https://www.youtube.com/watch?v=xKrIp4ZVOrg). Whole presentation is worth to hear for so many reasons.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [November 2, 2020, 4:53am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/11 "2020-11-02T04:53:37Z")

</div>

I read the v1.5 release highlights, but I didn’t think the allocation optimization would apply to how tuples would be stored. Just to clarify, am I correct in thinking that particular part of the video is showing that while v1.4 allocated a million tuples and an array that points to them, v1.5 allocated just the array that stored the tuples directly?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [November 2, 2020, 4:58am UTC](https://discourse.julialang.org/t/when-julia-gets-within-1-3x-of-c-c-speed-why-is-c-c-usually-faster/49341/12 "2020-11-02T04:58:47Z")

</div>

Exactly.
