# Are branches in Julia a problem for performance?

**URL:** https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062
**Category:** Performance
**Created:** [December 26, 2023, 5:18pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062 "2023-12-26T17:18:45Z")
**Posts on this page:** 11
**Page:** 1

<div class="post-metadata">

### Author: ![KZiemian](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kziemian/32/9020_2.png) [@KZiemian](https://discourse.julialang.org/u/KZiemian)
#### Post date: [December 26, 2023, 5:18pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/1 "2023-12-26T17:18:45Z")

</div>

Due to how modern CPU and computer memory works, branching, like using [`if` statement](https://sdremthix.medium.com/branchless-programming-why-your-cpu-will-thank-you-5f405d97b0c8), can produce high cost for the performance and produce [cache misses](https://redis.com/glossary/cache-miss/). At basic level I understand that, but I’m deficiently not expert in HPC.

More specific, I don’t know what is the cost of branching in Julia and when we should start even bother about such things. I quickly checked [Performance Tips](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-tips) and didn’t find relevant informations.

---

<div class="post-metadata">

### Author: ![Zentrik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zentrik/32/35409_2.png) [@Zentrik](https://discourse.julialang.org/u/Zentrik)
#### Post date: [December 26, 2023, 5:36pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/2 "2023-12-26T17:36:57Z")

</div>

The cost of branching in Julia is the same as in languages like C++; as the articles mention you can increase speed by not branching. However this is an optimization that generally will be on the order of nanoseconds and so is probably one of the last things you would do to improve performance.

---

<div class="post-metadata">

### Author: ![AMJ](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amj/32/214096_2.png) [@AMJ](https://discourse.julialang.org/u/AMJ)
#### Post date: [December 26, 2023, 7:15pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/3 "2023-12-26T19:15:31Z")

</div>

When you start to talk about branchless performance you are counting CPU cycles, i.e. L3 cache miss penalty is 50~70 CPU cycles or around 10 nanoseconds in modern CPUs. Before arriving at this level, you had to have thought about avoiding kernel space as much as possible by making the least amount of `syscall()`s(ranging from 20 nanoseconds to 300 nanoseconds on average). Now these are hardware-level/OS-specific protocols that are beyond a programming language’s scope. Julia’s performance tips will take you to 80% of the absolute performance, more than that you need to tinker with low-level stuff that is not that different than a C look-a-like code.

Edit: Also to add, usually when you are fetching data from L1 L2 cache repeatedly is worth the idea to go branchless, if you are fetching data from RAM your branchfull code will be faster most of the times.

---

<div class="post-metadata">

### Author: ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)
#### Post date: [December 27, 2023, 9:26am UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/4 "2023-12-27T09:26:59Z")

</div>

Yes, branches in Julia are a problem for performance (as they are in all high-performance programming languages). You can read more about how Julia and hardware interacts in [an old, but still relevant blogpost of mine](https://viralinstruction.com/posts/hardware/).

However, note that, like other compiled languages, Julia’s compiler will often remove branches where possible. For example, consider the step function of the Collatz Conjecture:

```julia
f(x) = isodd(x) ? 3x + 1 : x ÷ 2

```

Here, Julia’s LLVM backend will compile away the if/else statement to a `cmove` instruction, compiling the function to:

```julia
        push rbp
        mov rcx, rdi
        lea rax, [rdi + 2*rdi + 1]
        mov rbp, rsp
        shr rcx
        test dil, 1
        cmove rax, rcx
        pop rbp
        ret

```

Also, I don’t think branches can produce cache misses. They can produce branch mispredictions, which are also a source of CPU latency, but they are much less a delay than L3 cache misses.

Also also, modern CPU branch predictors are really really good, and I’ve often found that removing a branch in a tight loop doesn’t improve timing markedly, probably because it’s predicted with almost 100% accuracy.

---

<div class="post-metadata">

### Author: ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)
#### Post date: [December 27, 2023, 11:04am UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/5 "2023-12-27T11:04:08Z")

</div>

Cache misses can occur when the CPU mispredicts what memory to fetch. It won’t do anything that writes back to memory until it knows which branch is actually taken, but it will fetch memory and execute instructions speculatively and then just throw out the work if a predicted branch happens to have been wrong. This can lead to cache misses when the prediction is wrong.

---

<div class="post-metadata">

### Author: ![sob](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@sob](https://discourse.julialang.org/u/sob)
#### Post date: [December 27, 2023, 1:09pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/6 "2023-12-27T13:09:14Z")

</div>

For reference, to those who haven’t seen it yet:

> **[Infographics: Operation Costs in CPU Clock Cycles - IT Hare on Soft.ware](http://ithare.com/infographics-operation-costs-in-cpu-clock-cycles/)**
>
> Quote:"Back in 80s, it was possible to calculate the speed of the program just by looking at assembly."Another Quote:"keep in mind that these days compilers tend to ignore inline specifications more often than not"\[→\]

---

<div class="post-metadata">

### Author: ![KZiemian](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kziemian/32/9020_2.png) [@KZiemian](https://discourse.julialang.org/u/KZiemian)
#### Post date: [December 27, 2023, 2:10pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/7 "2023-12-27T14:10:33Z")

</div>

Fantastic! I will show it my students in C class! 😀

---

<div class="post-metadata">

### Author: ![sob](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@sob](https://discourse.julialang.org/u/sob)
#### Post date: [December 27, 2023, 6:36pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/8 "2023-12-27T18:36:24Z")

</div>

I first discovered it as it was linked on one of @ChrisRackauckas 's online courses.

---

<div class="post-metadata">

### Author: ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)
#### Post date: [December 27, 2023, 6:51pm UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/9 "2023-12-27T18:51:28Z")

</div>

My 2 cents worth. Look at CPU core pinning before you start to think about branches.

Ie. Keep those processes busy on the same core and don’t let the Is stall them and move to different cores

---

<div class="post-metadata">

### Author: ![sob](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@sob](https://discourse.julialang.org/u/sob)
#### Post date: [December 29, 2023, 8:18am UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/10 "2023-12-29T08:18:16Z")

</div>

And I’m not sure if this applies to Windows, but under Linux you can pass `isolcpus` as a kernel parameter to prevent some cores from being used by the scheduler without intervention. You could then manually force `julia` onto the excluded cores.

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [December 29, 2023, 8:55am UTC](https://discourse.julialang.org/t/are-branches-in-julia-a-problem-for-performance/108062/11 "2023-12-29T08:55:14Z")

</div>

Branchless can be faster but often isn’t faster at all because branch predictors can predict which way you’re gonna branch anyway. You can measure branch misses with `perf stat`. It’s good to measure them before trying to remove them.
