# Bring Intel x86 simd sort library to Julia

**URL:** <https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719>\
**Category:** General Usage\
**Tags:** package\
**Created:** [August 19, 2025, 8:00pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719 "2025-08-19T20:00:32Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:03pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/4 "2025-08-19T21:03:53Z")

</div>

I am not getting the same benchmark especially that i have avx-512 on my machine. Plus you are are not using multithreading. I have intel sort faster than julia sort by 6x and 20x with multithreading.  
Also look at your post:

> [@Numpy.sort vs Julia sort](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421/5):
>
> On a machine with AVX2 I can confirm numpy.sort is faster (but definitely not 10x): julia\> using BenchmarkTools, PyCall julia\> numpy = pyimport("numpy"); julia\> a = randn(Float32, 10^8); julia\> apy = pycall(numpy.array, PyObject, a); # convert to npy array julia\> @btime sort($a); @btime pycall($(numpy.sort), PyObject, $apy); 2.611 s (14 allocations: 850.76 MiB) 954.375 ms (1 allocation: 16 bytes) julia\> numpy.\_\_version\_\_ "2.1.3" julia\> versioninfo() Julia Version 1.11.2 Commit 5e9a32e…

I think you should compile with the -mavx512f ..etc flags, you also don’t need to build the library, you can use the templates in src.

---

_[View the full topic](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719)._
