# Bring Intel x86 simd sort library to Julia

**URL:** <https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719>\
**Category:** General Usage\
**Tags:** package\
**Created:** [August 19, 2025, 8:00pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719 "2025-08-19T20:00:32Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 8:00pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/1 "2025-08-19T20:00:32Z")

</div>

Hi,

Intel x86 simd sort library can be used via C++ template.

If I wanted to make this library available to the maximum of Julia users (a bit like numpy did) what would be the best steps to follow?

Should I use BinaryBuilder.jl?

Thank you

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [August 19, 2025, 8:22pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/2 "2025-08-19T20:22:19Z")

</div>

BinaryBuilder would be the way to do this. that said, as an alternative, it shouldn’t be too hard to copy the algorithm to Julia

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 8:54pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/3 "2025-08-19T20:54:17Z")

</div>

I’d start by checking whether that’s worth at all

```cxx
#include "x86simdsort.h"

extern "C" void qsort_float(float *arr, size_t size) {
    x86simdsort::qsort(arr, size, true);
}

extern "C" void qsort_double(double *arr, size_t size) {
    x86simdsort::qsort(arr, size, true);
}

```

Compiled with

```sh
g++ -o libsort.so -Wall -O3 -march=native -shared sort.c -L ../builddir -lx86simdsortcpp -Wl,-rpath,../builddir

```

Then

```julia-auto
julia> using BenchmarkTools, Random

julia> qsort!(x::Vector{Cfloat}) = @ccall "./libsort.so".qsort_float(x::Ptr{Cfloat}, length(x)::Csize_t)::Cvoid
qsort! (generic function with 1 method)

julia> qsort!(x::Vector{Cdouble}) = @ccall "./libsort.so".qsort_double(x::Ptr{Cdouble}, length(x)::Csize_t)::Cvoid
qsort! (generic function with 2 methods)

julia> @benchmark qsort!(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 140 samples with 1 evaluation per sample.
 Range (min … max): 26.242 ms … 38.469 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 36.768 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 34.580 ms ± 4.452 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

  █ ▁ ▁▃ ▂ ▁ ▂
  █▇▃▁▃▁▁▁▁▁▃▁▁▃▁▁▁▁▁▁▁▁▁▁▁▃▁▁▁▃▃▁▁▁▁▁▁▃▁▁▁▁▃▁▁▄▃▅█▅██▇███▆█▄ ▃
  26.2 ms Histogram: frequency by time 38.4 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sort!(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 602 samples with 1 evaluation per sample.
 Range (min … max): 4.731 ms … 11.296 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 6.740 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 6.880 ms ± 1.054 ms ┊ GC (mean ± σ): 0.05% ± 0.62%

    ▃ ▅ █▄ ▁▂▄
  ▃▇█▄▁▃▁▁▁▁▂▁▃▃▅▆█▆▄▇▇███▆▅▄▃▄▃▄▅▆▇█████▄▁▁▁▃▁▁▁▁▁▂▁▁▁▂▁▁▁▃ ▄
  4.73 ms Histogram: frequency by time 9.79 ms <

 Memory estimate: 4.01 MiB, allocs estimate: 6.

julia> @benchmark qsort!(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 88 samples with 1 evaluation per sample.
 Range (min … max): 44.550 ms … 66.835 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 55.813 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 54.247 ms ± 4.553 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

                                        ▆██▅▂ ▃
  ▄▄██▄▄▄▁▁▄▅▁▁▁▁▁▄▁▁▄▁▄▁▁▁▄▄▁▅▁▁▁▄▁▄▄▄▅█████▇█▄▅▄▅▄▁▄▁▄▄▁▁▁▄ ▁
  44.5 ms Histogram: frequency by time 61.3 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sort!(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 308 samples with 1 evaluation per sample.
 Range (min … max): 9.046 ms … 17.033 ms ┊ GC (min … max): 1.97% … 1.97%
 Time (median): 14.814 ms ┊ GC (median): 1.94%
 Time (mean ± σ): 14.138 ms ± 2.139 ms ┊ GC (mean ± σ): 4.03% ± 5.14%

                         ▂█ ▃ ▄▄▁▄▁
  ▃▁▃▅▆█▃▃▁▁▁▁▁▁▁▁▃▃▃▃▄▄▆███▇▆▃▄▃▁▃▃▁▄▃▇▅▃▅▄▆▅▆▅▆▆▆▇█▇█████▇▄ ▄
  9.05 ms Histogram: frequency by time 16.9 ms <

 Memory estimate: 8.01 MiB, allocs estimate: 6.

```

Tested on

```julia-auto
julia> versioninfo()
Julia Version 1.11.6
Commit 9615af0f269 (2025-07-09 12:58 UTC)
Build Info:
  Official https://julialang.org/ release
Platform Info:
  OS: Linux (x86_64-linux-gnu)
  CPU: 192 × AMD Ryzen Threadripper PRO 7995WX 96-Cores
  WORD_SIZE: 64
  LLVM: libLLVM-16.0.6 (ORCJIT, znver4)
Threads: 1 default, 0 interactive, 1 GC (on 192 virtual cores)

```

To be clear, sorting algorithms are different, but what do you want to achieve?

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:03pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/4 "2025-08-19T21:03:53Z")

</div>

I am not getting the same benchmark especially that i have avx-512 on my machine. Plus you are are not using multithreading. I have intel sort faster than julia sort by 6x and 20x with multithreading.  
Also look at your post:

> [@Numpy.sort vs Julia sort](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421/5):
>
> On a machine with AVX2 I can confirm numpy.sort is faster (but definitely not 10x): julia\> using BenchmarkTools, PyCall julia\> numpy = pyimport("numpy"); julia\> a = randn(Float32, 10^8); julia\> apy = pycall(numpy.array, PyObject, a); # convert to npy array julia\> @btime sort($a); @btime pycall($(numpy.sort), PyObject, $apy); 2.611 s (14 allocations: 850.76 MiB) 954.375 ms (1 allocation: 16 bytes) julia\> numpy.\_\_version\_\_ "2.1.3" julia\> versioninfo() Julia Version 1.11.2 Commit 5e9a32e…

I think you should compile with the -mavx512f ..etc flags, you also don’t need to build the library, you can use the templates in src.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 9:07pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/5 "2025-08-19T21:07:09Z")

</div>

> [@gitboy16](#):
>
> I am not getting the same benchmark especially that i have avx-512 on my machine.

Care to share what you tried?

> [@gitboy16](#):
>
> Plus you are are not using multithreading.

Of course I didn’t, julia’s `sort!` isn’t multi-threaded, so the comparison would be unfair.

> [@gitboy16](#):
>
> I 5hink you should compile with the -mavx512f ..etc flags,

Did you notice the `-march=native`? Also, did you read the [README of the library](https://github.com/intel/x86-simd-sort/blob/058f9132b87bb11ab7dd1aa27e07b070c6ba0f4b/README.md#x86-simd-sort)?

> The library auto picks the best version depending on the processor it is run on.

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:12pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/6 "2025-08-19T21:12:22Z")

</div>

Read this instead that is what I tried and with avx-512 julia sort is no way near it. It completely left in the dust for a vector of 10^8 doubles.

> <https://github.com/intel/x86-simd-sort/blob/main/src/README.md>

It is much easier to reproduce assuming you have the required hardware.

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:14pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/7 "2025-08-19T21:14:40Z")

</div>

If you don’t have a intel processor with avx-512, there is no point comparing them. That is what the intel’ algorithm has been dsigned for.

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:19pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/8 "2025-08-19T21:19:34Z")

</div>

Thank you. I wish I could! But I don’t have the knowledge.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 19, 2025, 9:19pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/9 "2025-08-19T21:19:49Z")

</div>

> [@gitboy16](#):
>
> It completely left in the dust for a vector of 10^8 doubles.

> [@gitboy16](#):
>
> I have intel sort faster than julia sort by 6x and 20x with multithreading.

what?.. I think if it’s 20x faster single-thread, I might call julia left in the dust.

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:21pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/10 "2025-08-19T21:21:56Z")

</div>

20x using avx512 and multithreaded. 6x with avx512 only (and one thread)  
Julia sort is unfortunetely not implemented for multithreaded.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 9:24pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/11 "2025-08-19T21:24:16Z")

</div>

You insist on not showing the code you run, I insist on showing mine:

```cxx
#include "src/x86simdsort-static-incl.h"

extern "C" void qsort_float(float *arr, size_t size) {
    x86simdsortStatic::qsort(arr, size, true);
}

extern "C" void qsort_double(double *arr, size_t size) {
    x86simdsortStatic::qsort(arr, size, true);
}

```

Compiled with

```sh
g++ -o libsort.so -Wall -O3 -fPIC -march=native -shared sort.c

```

Then in Julia

```julia-auto
julia> using BenchmarkTools, Random

julia> qsort!(x::Vector{Cfloat}) = @ccall "./libsort.so".qsort_float(x::Ptr{Cfloat}, length(x)::Csize_t)::Cvoid
qsort! (generic function with 1 method)

julia> qsort!(x::Vector{Cdouble}) = @ccall "./libsort.so".qsort_double(x::Ptr{Cdouble}, length(x)::Csize_t)::Cvoid
qsort! (generic function with 2 methods)

julia> @benchmark qsort!(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 142 samples with 1 evaluation per sample.
 Range (min … max): 25.982 ms … 38.337 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 36.373 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 34.133 ms ± 4.446 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▆▃ ▃▂▄▁█▄▂▁
  ██▆▄▁▁▁▁▃▁▃▃▁▁▅▁▁▁▁▁▁▁▄▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▄▆▇████████▆▆▄ ▃
  26 ms Histogram: frequency by time 38.2 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sort!(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 625 samples with 1 evaluation per sample.
 Range (min … max): 4.554 ms … 9.635 ms ┊ GC (min … max): 0.00% … 1.84%
 Time (median): 6.556 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 6.653 ms ± 980.903 μs ┊ GC (mean ± σ): 0.32% ± 1.70%

  ▁▆ █ ▂▃▄▁▆▃ ▂▂▂▅▃▆
  ██▃▃▁▁▁▁▁▁▁▁▁▂▁▅▅▆██▅▇▇▇▅██████▇▆▄▄▃▆▃▅▅▆▅▇███████▆▄▃▂▁▁▃▂▂ ▄
  4.55 ms Histogram: frequency by time 8.51 ms <

 Memory estimate: 4.01 MiB, allocs estimate: 6.

julia> @benchmark qsort!(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 83 samples with 1 evaluation per sample.
 Range (min … max): 44.114 ms … 63.417 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 60.562 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 57.460 ms ± 6.464 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

                                                    ▁█▄▄▇
  ▅▆▅▅█▅▁▃▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▁▃▁▁▁▁▁▁▁▁▁▁▅▁▃▁▁▃▃▁▁█▆██████▅▆█ ▁
  44.1 ms Histogram: frequency by time 63 ms <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> @benchmark sort!(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 303 samples with 1 evaluation per sample.
 Range (min … max): 9.049 ms … 26.193 ms ┊ GC (min … max): 1.96% … 0.49%
 Time (median): 14.754 ms ┊ GC (median): 1.53%
 Time (mean ± σ): 14.044 ms ± 2.186 ms ┊ GC (mean ± σ): 3.04% ± 4.35%

                      ▁▃▄ ▇ ▂ ▁▁ ▁▆█▅█▃▄▅▆▃▁
  ▃▅▅▅▆▃▄▄▆▅▃▃▁▁▁▁▃▃▅▄█████▅█▅█▄▇▅▄▄▄▄▃█▄▃█▅███▇███████████▆▆ ▅
  9.05 ms Histogram: frequency by time 16.8 ms <

 Memory estimate: 8.01 MiB, allocs estimate: 6.

```

which, unsurprisingly, matches the benchmarks I did above.

> [@gitboy16](#):
>
> If you don’t have a intel processor with avx-512, there is no point comparing them.

Did you bother checking the specs of the CPU I used, which [I had shared above](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/3)?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 19, 2025, 9:26pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/12 "2025-08-19T21:26:27Z")

</div>

> [@giordano](#):
>
> Did you bother checking the specs of the CPU I used, which [I had shared above](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/3)?

I think their point is it’s not an Intel CPU… which, idk, maybe Intel is doing the MKL shit again?

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 9:28pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/13 "2025-08-19T21:28:11Z")

</div>

> [@jling](#):
>
> I think their point is it’s not an Intel CPU… which, idk, maybe Intel is doing the MKL shit again?

I don’t think they can do that in open source code, [they just use intrinsics](https://github.com/intel/x86-simd-sort/blob/058f9132b87bb11ab7dd1aa27e07b070c6ba0f4b/src/avx2-64bit-qsort.hpp).

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:29pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/14 "2025-08-19T21:29:08Z")

</div>

Just use numpy, it is using it under the hood. I am on my phone, don’t have the code here. Plus it is c++ but i can post it later. And I said 10^8 double not 2^20 float. I will post everything later.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 9:36pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/15 "2025-08-19T21:36:41Z")

</div>

> [@gitboy16](#):
>
> Just use numpy

As a matter of fact, I did, and again got similar results.

```julia-auto
julia> using CondaPkg, PythonCall, BenchmarkTools, Random
[...]

(jl_paD8yw) pkg> conda add numpy
[...]

julia> numpy = pyimport("numpy")
Python: <module 'numpy' from '/tmp/jl_paD8yw/.CondaPkg/.pixi/envs/default/lib/python3.13/site-packages/numpy/ __init__.py'>

julia> @benchmark numpy.sort(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 139 samples with 1 evaluation per sample.
 Range (min … max): 25.406 ms … 46.719 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 35.977 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 35.103 ms ± 5.572 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▁ ▁▄█▂▃
  █▇██▅▁▃▃▅▁▃▁▁▁▁▃▁▁▁▃▁▁▁▁▁▁▁▅█████▇▆▇█▃▆▃▁▁▁▄▃▁▁▁▄▃▄▃▆▃▄▁▃▃▅ ▃
  25.4 ms Histogram: frequency by time 45.6 ms <

 Memory estimate: 632 bytes, allocs estimate: 29.

julia> @benchmark sort(x) setup=(Random.seed!(123); x = rand(Float32, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 578 samples with 1 evaluation per sample.
 Range (min … max): 4.619 ms … 10.924 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 7.478 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 7.329 ms ± 1.471 ms ┊ GC (mean ± σ): 2.69% ± 5.54%

   █ █ ▁▄▂▅▃▃▂▂
  ▆█▃▃▁▃▂▃▃▅██▄▃▅▆▅▄▄▅▄▇▆▇█████████▆▄▇▄▇▇▇▅▃▃▁▂▃▂▃▁▃▂▂▃▃▄▅▅▅ ▄
  4.62 ms Histogram: frequency by time 10.7 ms <

 Memory estimate: 8.01 MiB, allocs estimate: 9.

julia> @benchmark numpy.sort(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 82 samples with 1 evaluation per sample.
 Range (min … max): 44.839 ms … 71.773 ms ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 61.541 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 58.657 ms ± 6.979 ms ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▆ ▂▂▆▂▃▆█▃
  █▇▇▁▁▁▁▄▄▁▄▁▁▁▁▁▄▁▁▁▁▁▁▁▄▁▁▁▄▁▁▄▁▁▄▁▇████████▅▁▄▁▁▁▁▁▁▁▁▁▁▄ ▁
  44.8 ms Histogram: frequency by time 69.6 ms <

 Memory estimate: 616 bytes, allocs estimate: 28.

julia> @benchmark sort(x) setup=(Random.seed!(123); x = rand(Float64, 2 ^ 20)) evals=1
BenchmarkTools.Trial: 257 samples with 1 evaluation per sample.
 Range (min … max): 9.808 ms … 21.534 ms ┊ GC (min … max): 3.28% … 2.26%
 Time (median): 16.556 ms ┊ GC (median): 2.58%
 Time (mean ± σ): 16.506 ms ± 3.351 ms ┊ GC (mean ± σ): 6.49% ± 7.01%

                █ ▇▁ ▁ ▁▁▁
  ▃▃▁▁▁▁▁▁▁▁▃▅▄▆█▇██▃▃▃▁▁▁▁▁▁▁▃▃▄▅▅▇▆█▄▅▆▄▃▁▃▁▃▁▃▁▃▃▅▅██████▇ ▃
  9.81 ms Histogram: frequency by time 21.1 ms <

 Memory estimate: 16.01 MiB, allocs estimate: 9.

```

As I said above, these are just the same results as when using the library directly.

> [@gitboy16](#):
>
> And I said 10^8 double not 2^20 float.

```julia-auto
julia> @benchmark numpy.sort(x) setup=(Random.seed!(123); x = rand(Float64, 10 ^ 8)) evals=1
BenchmarkTools.Trial: 1 sample with 1 evaluation per sample.
 Single result which took 8.060 s (0.00% GC) to evaluate,
 with a memory estimate of 616 bytes, over 28 allocations.

julia> @benchmark sort(x) setup=(Random.seed!(123); x = rand(Float64, 10 ^ 8)) evals=1
BenchmarkTools.Trial: 2 samples with 1 evaluation per sample.
 Range (min … max): 2.516 s … 2.703 s ┊ GC (min … max): 0.05% … 5.43%
 Time (median): 2.609 s ┊ GC (median): 2.84%
 Time (mean ± σ): 2.609 s ± 132.008 ms ┊ GC (mean ± σ): 2.84% ± 3.81%

  █ █
  █▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ ▁
  2.52 s Histogram: frequency by time 2.7 s <

 Memory estimate: 1.49 GiB, allocs estimate: 9.

```

Anything else I should try?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 19, 2025, 9:37pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/16 "2025-08-19T21:37:40Z")

</div>

> [@gitboy16](#):
>
> And I said 10^8 double

with that much memory movement, are you sure you’re not seeing some GC artifact? i.e. Python turns off GC when you time it, Julia maybe not?

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:38pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/17 "2025-08-19T21:38:26Z")

</div>

We are clearly not getting the same results. Sorry for that.  
What I don’t understand is that in anotger thread that I linked above you are showing numpy outperforming julia. So it seems that you have inconsistent results.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [August 19, 2025, 9:44pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/18 "2025-08-19T21:44:00Z")

</div>

> [@gitboy16](#):
>
> So it seems that you have inconsistent results.

Did you notice the different CPU (that was AVX2, this is AVX512)?

---

<div class="post-metadata">

**Author:** ![gitboy16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gitboy16/32/24906_2.png) [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Post date:** [August 19, 2025, 9:58pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/19 "2025-08-19T21:58:43Z")

</div>

Which makes no sense at all 🤣  
Anyway let’s put that to bed, I think we are both wasting our time on this one.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [August 19, 2025, 10:30pm UTC](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719/20 "2025-08-19T22:30:55Z")

</div>

> [@gitboy16](#):
>
> If you don’t have a intel processor

Is the library really specific to Intel processors, not x86 with AVX2 or AVX512 in general?

[Next page](https://discourse.julialang.org/t/bring-intel-x86-simd-sort-library-to-julia/131719.md?page=2)
