# SIMD: Need some help to speed up sampling a code vector

**URL:** <https://discourse.julialang.org/t/simd-need-some-help-to-speed-up-sampling-a-code-vector/117075>\
**Category:** Performance\
**Tags:** sampling\
**Created:** [July 15, 2024, 9:39pm UTC](https://discourse.julialang.org/t/simd-need-some-help-to-speed-up-sampling-a-code-vector/117075 "2024-07-15T21:39:08Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![minetest2048](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/minetest2048/32/45961_2.png) [@minetest2048](https://discourse.julialang.org/u/minetest2048)\
**Post date:** [July 16, 2024, 3:46am UTC](https://discourse.julialang.org/t/simd-need-some-help-to-speed-up-sampling-a-code-vector/117075/3 "2024-07-16T03:46:40Z")

</div>

I like the fixed point number implementation better, its how hardware NCOs work on hardware GPS receiver

One option is to vectorize the fixed point index calculation and table lookup using gather instructions:

- [A SIMD intrinsic correlator library for GNSS software receivers | GPS Solutions](https://link.springer.com/article/10.1007/s10291-019-0865-8)

We can do vectorized load from a lookup table if we can somehow convince Julia to emit `vgatherdpd` instruction: [vgatherdps](https://www.officedaytime.com/simd512e/simdimg/si.php?f=vgatherdps) . Whether this is faster than a indexing loop on a CPU is debatable:

- [https://dl.acm.org/doi/abs/10.1145/3533737.3535089](https://dl.acm.org/doi/abs/10.1145/3533737.3535089)
- [SIMD gather result in slow down](https://discourse.julialang.org/t/simd-gather-result-in-slow-down/95161)
- [\> If you count shared memory scatter/gather, CPU SIMD already have both. Scatter... | Hacker News](https://news.ycombinator.com/item?id=19239273)
- [x86 - Intel vs AMD gather AVX performance - Stack Overflow](https://stackoverflow.com/questions/75845054/intel-vs-amd-gather-avx-performance)

The paper shows that its profitable for an i9-7900X processor with AVX512:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/c/1/c13634c0a24fdee5f5247c55f5a616aab575378e.png)

(`reg_standalone` is scalar indexing loop)

But a security update might make this fast vectorized lookup table code go 50% slower:

- [Compilation options for Downfall mitigation](https://discourse.julialang.org/t/compilation-options-for-downfall-mitigation/104844)
- [https://downfall.page/](https://downfall.page/)

Another option is to run the code LFSRs in parallel instead of indexing into a lookup table:

- [Generating more than one bit at a time with an LFSR](https://zipcpu.com/dsp/2017/11/13/lfsr-multi.html)
- [https://ufdcimages.uflib.ufl.edu/AA/00/03/94/72/00001/AA00039472\_00001.pdf](https://ufdcimages.uflib.ufl.edu/AA/00/03/94/72/00001/AA00039472_00001.pdf)

I haven’t seen anyone doing this for GNSS PRN generators though

---

_[View the full topic](https://discourse.julialang.org/t/simd-need-some-help-to-speed-up-sampling-a-code-vector/117075)._
