# Nbabel nbody integrator speed up

**URL:** <https://discourse.julialang.org/t/nbabel-nbody-integrator-speed-up/51712>\
**Category:** Performance\
**Tags:** question, review\
**Created:** [December 12, 2020, 11:26am UTC](https://discourse.julialang.org/t/nbabel-nbody-integrator-speed-up/51712 "2020-12-12T11:26:19Z")\
**Posts on this page:** 1\
**Showing post:** 22

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [December 12, 2020, 7:10pm UTC](https://discourse.julialang.org/t/nbabel-nbody-integrator-speed-up/51712/22 "2020-12-12T19:10:33Z")

</div>

> [@baggepinnen](#):
>
> I see, that’s very good to know! Does this reasoning extend to `SVector{3}` as well, which I assume is packed into a vector like a matrix of 3 x N? I could then potentially slow down my code by switching from the N x 3 matrix layout to a vector of `SVector` ?

You could. [HybridArrays](https://github.com/mateuszbaran/HybridArrays.jl) may give you the best of both worlds, by making them N x 3. A slice `A[n,:]` should still return an `SVector`. If everything inlines, it should still be able to SIMD across loop iterations, while giving you the convenience of expressing some operations on the vectors instead making everything loops.

> [@baggepinnen](#):
>
> I did a similar benchmark to yours, and found that the N x 4 layout was faster than the 4 x N as well, but 4 should be a good value to SIMD on? I had to go up to around N x 16 before what I previously thought was the better memory layout to actually be faster. Is it so hard for the compiler to apply SIMD within a single loop iteration that the arrays have to be large enough for cache locality to become important before the column-major access wins out?

Another factor to consider is unrolling. In the example with N x 4 vs 4 x N, the N x 4 case will SIMD and be 4x unrolled (one per column). The 4 x N will only do a single operation per loop iteration.

Theoretically, when you don’t have dependencies (i.e, `s += x[i]`, where each iteration depends on the previous), your CPU should be able to execute different loop iterations in parallel via out of order processing + speculative execution, but in practice, I normally find some unrolling tends to help.  
Maybe it’s because of better out of order, or maybe it’s because of better density of relevant instructions, vs things like incrementing and checking loop counters.

---

_[View the full topic](https://discourse.julialang.org/t/nbabel-nbody-integrator-speed-up/51712)._
