# \#simd

**URL:** https://discourse.julialang.org/tag/simd/74.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [Performance challenge: can you write a faster sum?](https://discourse.julialang.org/t/performance-challenge-can-you-write-a-faster-sum/130456)

<div class="topic-metadata">

**Author:** [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Replies:** 32\
**Last updated:** [December 11, 2025, 10:03pm UTC](https://discourse.julialang.org/t/performance-challenge-can-you-write-a-faster-sum/130456 "2025-12-11T22:03:55Z")

</div>

OK, I’ve been fiddling with this on and off for a month now, but I’m curious if anyone can do better. Here’s the challenge: write a single generic implementation of sum(A) such that it’s fast for both A = rand(10000) an…

---

## [Multithreading non-contiguous mapreduction over huge matrices](https://discourse.julialang.org/t/multithreading-non-contiguous-mapreduction-over-huge-matrices/134413)

<div class="topic-metadata">

**Author:** [@noetheriankoala](https://discourse.julialang.org/u/noetheriankoala)\
**Replies:** 5\
**Last updated:** [December 8, 2025, 7:36am UTC](https://discourse.julialang.org/t/multithreading-non-contiguous-mapreduction-over-huge-matrices/134413 "2025-12-08T07:36:45Z")

</div>

Hello! A version of the following function is majorly bottle-necking a project I am working on and I could use help speeding it up. In all the below code, A is an k x n matrix, l is a length-n vector, and q is a length n…

---

## [Vectorizing an awkward computation: any Julia magic?](https://discourse.julialang.org/t/vectorizing-an-awkward-computation-any-julia-magic/133212)

<div class="topic-metadata">

**Author:** [@moble](https://discourse.julialang.org/u/moble)\
**Replies:** 6\
**Last updated:** [October 19, 2025, 4:19am UTC](https://discourse.julialang.org/t/vectorizing-an-awkward-computation-any-julia-magic/133212 "2025-10-19T04:19:06Z")

</div>

I’m computing a matrix of values (specifically, a Wigner D matrix for a given ℓ value) via recurrence for a given input value R. The nature of this recurrence makes the memory-access patterns pretty ugly, and there’s re…

---

## [SIMD.jl - vload() continuous blocks from higher-dimensional arrays?](https://discourse.julialang.org/t/simd-jl-vload-continuous-blocks-from-higher-dimensional-arrays/131981)

<div class="topic-metadata">

**Author:** [@jl\_enthusiast](https://discourse.julialang.org/u/jl_enthusiast)\
**Replies:** 4\
**Last updated:** [August 31, 2025, 9:45pm UTC](https://discourse.julialang.org/t/simd-jl-vload-continuous-blocks-from-higher-dimensional-arrays/131981 "2025-08-31T21:45:29Z")

</div>

Hi, I am currently doing some practice projects in Julia, which involves using SIMD.jl for writing/generating microkernels. My issue is trying to determine what is the “best” way of loading Vec types from a higher-dimen…

---

## [Huge performance variance with \`if\` options in loops](https://discourse.julialang.org/t/huge-performance-variance-with-if-options-in-loops/130462)

<div class="topic-metadata">

**Author:** [@hz-xiaxz](https://discourse.julialang.org/u/hz-xiaxz)\
**Replies:** 4\
**Last updated:** [July 4, 2025, 6:51am UTC](https://discourse.julialang.org/t/huge-performance-variance-with-if-options-in-loops/130462 "2025-07-04T06:51:35Z")

</div>

I’m trying to optimize a non-allocating matrix update. Here is the simple code. function update\_W\_original!( W::AbstractMatrix, l::Int, K::Int, col\_cache::AbstractVector, row\_cache::AbstractVector, )…

---

## [How to make the most of SIMD.jl when number of data elements is not divisible by SIMD width](https://discourse.julialang.org/t/how-to-make-the-most-of-simd-jl-when-number-of-data-elements-is-not-divisible-by-simd-width/129226)

<div class="topic-metadata">

**Author:** [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Replies:** 5\
**Last updated:** [May 22, 2025, 10:51am UTC](https://discourse.julialang.org/t/how-to-make-the-most-of-simd-jl-when-number-of-data-elements-is-not-divisible-by-simd-width/129226 "2025-05-22T10:51:36Z")

</div>

Consider the following MWE that simply stores in c an elementwise vector multiplication (so c=a.\*b). Here I want to take SIMD chunks and do the operation in SIMD vectors. But the “naive solution” that I get is around 2x…

---

## [Datatypes SIMD by default? (as in Mojo)](https://discourse.julialang.org/t/datatypes-simd-by-default-as-in-mojo/128777)

<div class="topic-metadata">

**Author:** [@alexandergunnarson](https://discourse.julialang.org/u/alexandergunnarson)\
**Replies:** 14\
**Last updated:** [May 15, 2025, 4:40pm UTC](https://discourse.julialang.org/t/datatypes-simd-by-default-as-in-mojo/128777 "2025-05-15T16:40:22Z")

</div>

Hello all! Super excited to use Julia. It won out in a lengthy evaluation of mine against Rust, “modern” C++ (\>= C++23), modern Java (\>= Java 24), Mojo, etc. as the basis of my startup’s backend. Thank you to all who hav…

---

## [Is it ok to iterate inside a \`@simd\` loop?](https://discourse.julialang.org/t/is-it-ok-to-iterate-inside-a-simd-loop/128338)

<div class="topic-metadata">

**Author:** [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Replies:** 8\
**Last updated:** [April 25, 2025, 10:29am UTC](https://discourse.julialang.org/t/is-it-ok-to-iterate-inside-a-simd-loop/128338 "2025-04-25T10:29:35Z")

</div>

I’m delightfully surprised that this works: function mr(f, op, A, n) a1, s = iterate(A) a2, s = iterate(A, s) v = op(f(a1), f(a2)) @simd for \_ in 3:n ai, s = iterate(A, s) v = op(v, f(ai)…

---

## [Min/max swap using SIMD](https://discourse.julialang.org/t/min-max-swap-using-simd/126832)

<div class="topic-metadata">

**Author:** [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Replies:** 18\
**Last updated:** [March 12, 2025, 11:35am UTC](https://discourse.julialang.org/t/min-max-swap-using-simd/126832 "2025-03-12T11:35:22Z")

</div>

I would like to request some help to SIMD the following code: import Test: @test n = 10\_000\_000 A = repeat(\[3, 1, 4, 2, 7, 8, 5, 6\], outer=n); B = repeat(\[5, 6, 0, 7, 2, 3, 1, 4\], outer=n); C = copy(A); D = copy(B); …

---

## [Autovectorization in Julia 101](https://discourse.julialang.org/t/autovectorization-in-julia-101/123472)

<div class="topic-metadata">

**Author:** [@Benny](https://discourse.julialang.org/u/Benny)\
**Replies:** 2\
**Last updated:** [December 5, 2024, 11:43am UTC](https://discourse.julialang.org/t/autovectorization-in-julia-101/123472 "2024-12-05T11:43:22Z")

</div>

I’m trying to read up on autovectorization but I might be looking in the wrong place: Auto-Vectorization in LLVM — LLVM 20.0.0git documentation For one, the whole thing talks about Clang and C code, so I have doubts it…

---

## [Experiments with LoopVectorization and convolutions](https://discourse.julialang.org/t/experiments-with-loopvectorization-and-convolutions/123188)

<div class="topic-metadata">

**Author:** [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Replies:** 24\
**Last updated:** [December 3, 2024, 8:35pm UTC](https://discourse.julialang.org/t/experiments-with-loopvectorization-and-convolutions/123188 "2024-12-03T20:35:27Z")

</div>

I have been doing some experiments with the great LoopVectorization for convolution, and I must admit that in spite of my efforts to manually refactor my code, I have been unable to even approach the speed that this pack…

---

## [Optimizing sums of products (dot products)](https://discourse.julialang.org/t/optimizing-sums-of-products-dot-products/119515)

<div class="topic-metadata">

**Author:** [@s-baumann](https://discourse.julialang.org/u/s-baumann)\
**Replies:** 17\
**Last updated:** [September 24, 2024, 7:40pm UTC](https://discourse.julialang.org/t/optimizing-sums-of-products-dot-products/119515 "2024-09-24T19:40:11Z")

</div>

Is there a way to avoid the intermediate allocation for a sumproduct scenario? In the following case I can redo it with a for loop which improves the memory allocation but causes it to take longer to execute (the differ…

---

## [Vectorization of multivariate normal PDF](https://discourse.julialang.org/t/vectorization-of-multivariate-normal-pdf/116248)

<div class="topic-metadata">

**Author:** [@xor0110](https://discourse.julialang.org/u/xor0110)\
**Replies:** 1\
**Last updated:** [June 26, 2024, 7:57am UTC](https://discourse.julialang.org/t/vectorization-of-multivariate-normal-pdf/116248 "2024-06-26T07:57:54Z")

</div>

I would like to compute the PDF of a multivariate normal distribution of given parameters for a number of input points. (eg MvNorm for several vectors at once) This is for a particle filter, many such computations going …

---

## [Failure to vectorize 8 Int64 multiplies when 8 Float64 multiplies vectorize](https://discourse.julialang.org/t/failure-to-vectorize-8-int64-multiplies-when-8-float64-multiplies-vectorize/114939)

<div class="topic-metadata">

**Author:** [@brainandforce](https://discourse.julialang.org/u/brainandforce)\
**Replies:** 8\
**Last updated:** [May 31, 2024, 11:52pm UTC](https://discourse.julialang.org/t/failure-to-vectorize-8-int64-multiplies-when-8-float64-multiplies-vectorize/114939 "2024-05-31T23:52:39Z")

</div>

I’ve come across this interesting observation in my package, CliffordNumbers.jl. For context, the multiplication done here is a geometric product (relevant code here) implemented as a grid multiply between blade coeffici…

---

## [SIMD.jl/shufflevector without support for SIMD vector as mask?](https://discourse.julialang.org/t/simd-jl-shufflevector-without-support-for-simd-vector-as-mask/110079)

<div class="topic-metadata">

**Author:** [@Julia2001](https://discourse.julialang.org/u/Julia2001)\
**Replies:** 2\
**Last updated:** [February 11, 2024, 9:19pm UTC](https://discourse.julialang.org/t/simd-jl-shufflevector-without-support-for-simd-vector-as-mask/110079 "2024-02-11T21:19:03Z")

</div>

Hi! Could anybody explain a bit on why SIMD.jl/shufflevector does not accept a Vec{N, T} object as “mask” argument. Wouldn’t it be a natural choice to allow “mask” to be a SIMD vector? Is this an LLVM limitation? Is …

---

## [Supporting SIMD-enabled objective functions in optimization APIs](https://discourse.julialang.org/t/supporting-simd-enabled-objective-functions-in-optimization-apis/109728)

<div class="topic-metadata">

**Author:** [@nielsls](https://discourse.julialang.org/u/nielsls)\
**Replies:** 3\
**Last updated:** [February 5, 2024, 8:22pm UTC](https://discourse.julialang.org/t/supporting-simd-enabled-objective-functions-in-optimization-apis/109728 "2024-02-05T20:22:32Z")

</div>

With SIMD the CPU may execute the same instructions on e.g. 4x differents datasets just as fast as on 1x dataset. Assuming the user has a SIMD-enabled objective function - are there optimization algos that can exploit t…

---

## [Different \`@code\_llvm\` output on macos and x86](https://discourse.julialang.org/t/different-code-llvm-output-on-macos-and-x86/107338)

<div class="topic-metadata">

**Author:** [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Replies:** 4\
**Last updated:** [December 8, 2023, 7:17pm UTC](https://discourse.julialang.org/t/different-code-llvm-output-on-macos-and-x86/107338 "2023-12-08T19:17:31Z")

</div>

Hi, I wonder about the different outputs I obtain from @code\_llvm with the same Julia version (1.10.rc2) on different architectures (arm vs x86). The following script: f(a,b) = a .+ b f (generic function with 1 method…

---

## [Optimizing Direct 2D Convolution Code](https://discourse.julialang.org/t/optimizing-direct-2d-convolution-code/106598)

<div class="topic-metadata">

**Author:** [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Replies:** 14\
**Last updated:** [November 23, 2023, 8:48pm UTC](https://discourse.julialang.org/t/optimizing-direct-2d-convolution-code/106598 "2023-11-23T20:48:40Z")

</div>

I have implemented 2D convolution using direct calculation: using BenchmarkTools; function \_Conv2D!( mO :: Matrix{T}, mI :: Matrix{T}, mK :: Matrix{T} ) where {T \<: AbstractFloat} numRowsI, numColsI = size(mI); …

---

## [Understanding the performance and overhead of a vector of SOA vs a vector of AOS for SIMD and the effect of push!](https://discourse.julialang.org/t/understanding-the-performance-and-overhead-of-a-vector-of-soa-vs-a-vector-of-aos-for-simd-and-the-effect-of-push/100560)

<div class="topic-metadata">

**Author:** [@f.ij](https://discourse.julialang.org/u/f.ij)\
**Replies:** 1\
**Last updated:** [June 23, 2023, 11:13am UTC](https://discourse.julialang.org/t/understanding-the-performance-and-overhead-of-a-vector-of-soa-vs-a-vector-of-aos-for-simd-and-the-effect-of-push/100560 "2023-06-23T11:13:23Z")

</div>

I’ve been working on an interactive simulation tool for Monte Carlo simulations of Ising(-like) Models. Performance is key, and something I’ve been a bit confused by. I made a very simplified version of the program here,…

---

## [Is this a valid use of simd?](https://discourse.julialang.org/t/is-this-a-valid-use-of-simd/100463)

<div class="topic-metadata">

**Author:** [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Replies:** 2\
**Last updated:** [June 16, 2023, 6:26pm UTC](https://discourse.julialang.org/t/is-this-a-valid-use-of-simd/100463 "2023-06-16T18:26:57Z")

</div>

I am trying to calculate the sum of the absolute differences between two vectors without allocation. I have no experience with @simd but was wondering if this is a correct use: function sum\_absdiff(x::Vector{T}, y::Vect…

---

## [Optimizing Direct 1D Convolution Code](https://discourse.julialang.org/t/optimizing-direct-1d-convolution-code/97658)

<div class="topic-metadata">

**Author:** [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Replies:** 21\
**Last updated:** [April 28, 2023, 2:53pm UTC](https://discourse.julialang.org/t/optimizing-direct-1d-convolution-code/97658 "2023-04-28T14:53:46Z")

</div>

I am trying to create a function for a direct (Not fft() based) 1D convolution. Originally I used the classic implementation: using BenchmarkTools; function \_Conv1D!( vO :: Array{T, 1}, vA :: Array{T, 1}, vB :: Array{…

---

## [\`\`\`@turbo\`\`\` producing different (and wrong) results compared to \`\`\`@inbounds @simd\`\`\`](https://discourse.julialang.org/t/turbo-producing-different-and-wrong-results-compared-to-inbounds-simd/96851)

<div class="topic-metadata">

**Author:** [@bremez](https://discourse.julialang.org/u/bremez)\
**Replies:** 3\
**Last updated:** [March 30, 2023, 5:34pm UTC](https://discourse.julialang.org/t/turbo-producing-different-and-wrong-results-compared-to-inbounds-simd/96851 "2023-03-30T17:34:30Z")

</div>

I have a use case where an integer array contains only values between 1 and some maxValue, and I would like the histogram of the array values. I found that counts(array, maxValue) was too slow, and as best as I could und…

---

## [Question on multithreading/vectorizing loops](https://discourse.julialang.org/t/question-on-multithreading-vectorizing-loops/95291)

<div class="topic-metadata">

**Author:** [@Ender\_L](https://discourse.julialang.org/u/Ender_L)\
**Replies:** 9\
**Last updated:** [March 22, 2023, 8:14pm UTC](https://discourse.julialang.org/t/question-on-multithreading-vectorizing-loops/95291 "2023-03-22T20:14:29Z")

</div>

Consider a classic loop parallelization scheme: the outer loop multithreads, and the inner loop vectorizes: tmp = \[Vector{T}(undef, length(B)) for \_ in 1:Threads.nthreads()\] @inbounds Threads.@threads for ii in…

---

## [Major performance boost when precaching random inputs to \`\`\`exp\`\`\`?](https://discourse.julialang.org/t/major-performance-boost-when-precaching-random-inputs-to-exp/85740)

<div class="topic-metadata">

**Author:** [@bremez](https://discourse.julialang.org/u/bremez)\
**Replies:** 8\
**Last updated:** [September 25, 2022, 1:28pm UTC](https://discourse.julialang.org/t/major-performance-boost-when-precaching-random-inputs-to-exp/85740 "2022-09-25T13:28:41Z")

</div>

I’m writing a Monte Carlo simulation which requires exponentials of uniformly distributed random numbers. Calling exp takes a sizeable portion of my my inner loop runtime, and so I thought to cache pre-calculated value. …

---

## [Vectorize but break early?](https://discourse.julialang.org/t/vectorize-but-break-early/87362)

<div class="topic-metadata">

**Author:** [@abulak](https://discourse.julialang.org/u/abulak)\
**Replies:** 3\
**Last updated:** [September 20, 2022, 3:10pm UTC](https://discourse.julialang.org/t/vectorize-but-break-early/87362 "2022-09-20T15:10:14Z")

</div>

I’d like to properly vectorize this: """ issuffix(u, v) Check if \`u\` is a suffix of \`v\` """ function issuffix(u::AbstractVector, v::AbstractVector) lu = length(u) lu ≤ length(v) || return false @inbounds…

---

## [LoopVectorization: @turbo performs worse than @inbounds on trivial loop](https://discourse.julialang.org/t/loopvectorization-turbo-performs-worse-than-inbounds-on-trivial-loop/66884)

<div class="topic-metadata">

**Author:** [@kevinsa5](https://discourse.julialang.org/u/kevinsa5)\
**Replies:** 9\
**Last updated:** [August 28, 2021, 5:50pm UTC](https://discourse.julialang.org/t/loopvectorization-turbo-performs-worse-than-inbounds-on-trivial-loop/66884 "2021-08-28T17:50:24Z")

</div>

I am a long-time admirer of julia who is finally getting serious about learning it. I am comparing the timings when applying @turbo, @inbounds, and @simd to the loop in a simple element-wise multiplication function and …

---

## [PaddedViews very slow](https://discourse.julialang.org/t/paddedviews-very-slow/66979)

<div class="topic-metadata">

**Author:** [@weymouth](https://discourse.julialang.org/u/weymouth)\
**Replies:** 7\
**Last updated:** [August 25, 2021, 3:07pm UTC](https://discourse.julialang.org/t/paddedviews-very-slow/66979 "2021-08-25T15:07:05Z")

</div>

I’m sure to be doing something silly, but I can’t seem to get PaddedViews.jl up to speed. Consider the LoopVectorization.jl image filtering example using LoopVectorization, OffsetArrays, Images, PaddedViews kern = Image…

---

## [Why is this @simd loop faster than a while loop even if it has longer assembly?](https://discourse.julialang.org/t/why-is-this-simd-loop-faster-than-a-while-loop-even-if-it-has-longer-assembly/65585)

<div class="topic-metadata">

**Author:** [@louie4825](https://discourse.julialang.org/u/louie4825)\
**Replies:** 6\
**Last updated:** [August 1, 2021, 7:17am UTC](https://discourse.julialang.org/t/why-is-this-simd-loop-faster-than-a-while-loop-even-if-it-has-longer-assembly/65585 "2021-08-01T07:17:25Z")

</div>

Why does @simd here make the loop run faster even if the while version has shorter generated assembly? (Thiscode is part of a prime sieve implementation.) using BenchmarkTools # Auxillary functions begin const \_uint\_b…

---

## [SIMD Complex Numbers](https://discourse.julialang.org/t/simd-complex-numbers/65038)

<div class="topic-metadata">

**Author:** [@danny.s](https://discourse.julialang.org/u/danny.s)\
**Replies:** 19\
**Last updated:** [July 22, 2021, 9:24pm UTC](https://discourse.julialang.org/t/simd-complex-numbers/65038 "2021-07-22T21:24:16Z")

</div>

Hi, I hope this hasn’t been discussed before, but has SIMD support for complex numbers been considered? There are several libraries in C/C++ taking advantage of this. There’s an example of implementation here and vector…

---

## [LoopVectorization.jl vmap gives an error ::VectorizationBase.Vec{4, Int64}](https://discourse.julialang.org/t/loopvectorization-jl-vmap-gives-an-error-vectorizationbase-vec-4-int64/65056)

<div class="topic-metadata">

**Author:** [@Storopoli](https://discourse.julialang.org/u/Storopoli)\
**Replies:** 17\
**Last updated:** [July 22, 2021, 8:12am UTC](https://discourse.julialang.org/t/loopvectorization-jl-vmap-gives-an-error-vectorizationbase-vec-4-int64/65056 "2021-07-22T08:12:57Z")

</div>

I am trying to do a vmap on a function. map and ThreadsX.map both work. collatz(x::Int64) = if iseven(x) x ÷ 2 else 3x + 1 end function collatz\_sequencia(x::Int64) n = 0 while true …

[Next page](https://discourse.julialang.org/tag/simd/74.md?match_all_tags=true&page=1&tags%5B%5D=simd)
