This is because you’re totally memory bound. Try @turbo with vectors of length 1024 or less.
Did you start julia with multiple threads?
EDIT:
Interestingly, I hadn’t actually seen LLVM SIMD min/max functions before, but it is now.
LV should still do better for sizes like 255, i.e. the power of 2-1.
Indeed you are correct. Comparing performance of myminmax2_basic, myminmax2_turbo and myminmax2_tturbo across different input lengths, @turbo wins for the smallest inputs, then in the range from se(12) to se(13)@tturbo takes the lead, after that myminmax2_basic is tied with myminmax2_turbo but myminmax2_tturbo is suddenly much slower (comparing the median timings):
julia> @benchmark myminmax2_basic(i) setup = (Random.seed!(12345678); i = se(14))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
Range (min … max): 9.457 μs … 43.592 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 9.759 μs ┊ GC (median): 0.00%
Time (mean ± σ): 10.439 μs ± 1.824 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▅█▆▅▄▄▃▃▃▃▃▃▂▂▂▂▂▂▂▂▁▁▁▁▁ ▁ ▂
████████████████████████████████▇▇▇▆▇▇▆▆▆▅▆▆▆▅▅▆▄▆▆▅▅▅▄▅▄▁▄ █
9.46 μs Histogram: log(frequency) by time 17.3 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
julia> @benchmark myminmax2_turbo(i) setup = (Random.seed!(12345678); i = se(14))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
Range (min … max): 9.418 μs … 93.085 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 9.709 μs ┊ GC (median): 0.00%
Time (mean ± σ): 10.176 μs ± 2.005 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▆█▆▅▅▄▄▃▂▂▂▂▁▁ ▂
█████████████████▇▇█▆▆▆▆▅▆▆▅▅▅▅▅▆▅▅▄▅▅▅▅▄▅▄▃▁▃▄▅▅▅▄▄▃▁▄▃▁▁▄ █
9.42 μs Histogram: log(frequency) by time 19.1 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
julia> @benchmark myminmax2_tturbo(i) setup = (Random.seed!(12345678); i = se(14))
BenchmarkTools.Trial: 9767 samples with 5 evaluations.
Range (min … max): 6.448 μs … 26.342 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 12.241 μs ┊ GC (median): 0.00%
Time (mean ± σ): 12.036 μs ± 1.755 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▂▂▃▅▆█▅▃▁
▁▁▁▁▂▂▃▃▂▂▂▂▁▂▁▁▂▂▂▃▃▄▅▆▆▅▆▆██████████▇▆▄▃▂▂▃▂▂▂▁▁▁▁▁▁▁▁▁▁▁ ▃
6.45 μs Histogram: frequency by time 16.9 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
julia> @benchmark myminmax2_basic(i) setup = (Random.seed!(12345678); i = se(15))
BenchmarkTools.Trial: 9742 samples with 1 evaluation.
Range (min … max): 18.815 μs … 119.003 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 23.825 μs ┊ GC (median): 0.00%
Time (mean ± σ): 24.777 μs ± 4.714 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▁▃▄▆█▇███▅▅▃▁▁
▄▅▅▇██████████████▇▆▅▅▄▃▃▃▃▃▂▂▂▂▂▂▂▂▂▂▂▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ ▃
18.8 μs Histogram: frequency by time 44.4 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
julia> @benchmark myminmax2_turbo(i) setup = (Random.seed!(12345678); i = se(15))
BenchmarkTools.Trial: 9707 samples with 1 evaluation.
Range (min … max): 18.815 μs … 106.610 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 23.704 μs ┊ GC (median): 0.00%
Time (mean ± σ): 24.672 μs ± 4.239 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▁▃▆▇█▇▆▅▄▃▁
▁▂▃▄▃▆████████████▆▆▅▄▄▄▃▃▂▂▂▂▂▂▂▂▂▂▁▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ ▃
18.8 μs Histogram: frequency by time 42 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
julia> @benchmark myminmax2_tturbo(i) setup = (Random.seed!(12345678); i = se(15))
BenchmarkTools.Trial: 5236 samples with 1 evaluation.
Range (min … max): 13.044 μs … 130.655 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 43.958 μs ┊ GC (median): 0.00%
Time (mean ± σ): 44.657 μs ± 8.647 μs ┊ GC (mean ± σ): 0.00% ± 0.00%
▂▃▅▆█▆▆▃▂
▁▁▁▁▁▁▁▁▂▂▂▂▂▂▁▂▃▃▂▂▂▂▃▃▅▇███████████▇▆▅▅▄▄▃▂▃▂▁▂▂▁▁▁▁▁▁▁▂▃▁ ▃
13 μs Histogram: frequency by time 72.2 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
This is with Julia started with four threads (same as above).
BTW, don’t think that I’m criticizing your work, I know that your packages are tremendously useful, it’s just that comparing microbenchmarks is something I enjoy very much . Also thank you for being so instructive here.