# Nerd-sniping: can you make this faster?

**URL:** <https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793>\
**Category:** Performance\
**Created:** [September 30, 2025, 8:10pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793 "2025-09-30T20:10:34Z")\
**Posts on this page:** 19\
**Page:** 2

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 1, 2025, 12:38pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/21 "2025-10-01T12:38:37Z")

</div>

I’ve pulled out the `deltaxy` calculation, so it could still be unfair there. But in practice (in the actual application), I get, with this last version:

```julia-repl
julia> @b atomic_sasa(prot; parallel=false)
13.640 ms (67044 allocs: 2.888 MiB)

```

while with Zentrik’s I get:

```julia-repl
julia> @b atomic_sasa(prot; parallel=false)
9.760 ms (67044 allocs: 2.888 MiB)

```

FWIW, this is what we gained so far:

```julia-auto
julia> @b atomic_sasa(prot; parallel=false)
80.444 ms (65546 allocs: 2.248 MiB)

```

(not irreleveant at all, since we need to run this calculation over thousands of frames of MD trajectories)

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [October 1, 2025, 12:42pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/22 "2025-10-01T12:42:08Z")

</div>

Not sure how much of a help it would be (if any), but

\sum\_{i=1}^n (\delta\_i + c\_i^j)^2 \> R^2, \forall j

can be also written as

\sum\_{i=1}^n (2\delta\_i + c\_i^j) c\_i^j \> R^2 - \sum\_{i=1}^n \delta\_i^2, \forall j.

The right-hand side does not depend on j, so it needs to be computed only once. I suppose the left-hand side would translate to

```julia
lhs = ((dot_cache.x[ind] + deltaxy[1]) * deltaxy[1] + 
       (dot_cache.y[ind] + deltaxy[2]) * deltaxy[2] + 
       (dot_cache.z[ind] + deltaxy[3]) * deltaxy[3])

```

and the right-hand side would be fixed to

```julia
rhs = rj_sq - sum(abs2, deltaxy)

```

Edit: And obviously the condition of interest:

```julia
exposed_i[ind] &= (lhs > rhs)

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 1, 2025, 12:58pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/23 "2025-10-01T12:58:25Z")

</div>

Nice catch, but there was no speedup:

```julia-auto
julia> @benchmark atomic_sasa($prot; parallel=false)
BenchmarkTools.Trial: 507 samples with 1 evaluation per sample.
 Range (min … max): 9.351 ms … 17.197 ms ┊ GC (min … max): 0.00% … 41.86%
 Time (median): 9.491 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 9.859 ms ± 820.727 μs ┊ GC (mean ± σ): 2.58% ± 6.11%

  ▃█▇▄▃▂▂▂▁▂ ▁           
  ██████████▇█▇▇▅▆██▇▇▇▇▄▅▅▄▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▅▅▅▇██▇▇▆▆▄▁▅▅ ▇
  9.35 ms Histogram: log(frequency) by time 12.2 ms <

 Memory estimate: 2.89 MiB, allocs estimate: 67044.

```

vs (previously):

```julia-auto
julia> @benchmark atomic_sasa($prot; parallel=false)
BenchmarkTools.Trial: 510 samples with 1 evaluation per sample.
 Range (min … max): 9.306 ms … 16.500 ms ┊ GC (min … max): 0.00% … 40.93%
 Time (median): 9.469 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 9.824 ms ± 829.032 μs ┊ GC (mean ± σ): 2.53% ± 6.20%

  ▁█▄                                                          
  ███▅▄▄▄▃▄▃▄▃▅▅▃▃▃▃▂▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂▁▁▁▁▁▁▁▁▁▂▂▂▂▃▂▃▃▂▂▁▂▃ ▃
  9.31 ms Histogram: frequency by time 12.2 ms <

 Memory estimate: 2.89 MiB, allocs estimate: 67044.

```

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [October 1, 2025, 1:30pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/24 "2025-10-01T13:30:22Z")

</div>

Btw, note that by using StructArrays.jl you can conveniently view the SoA also as an AoS.

```julia-auto
s = StructArray{SVector{3,Float32}}((dot_cache.x, dot_cache.y, dot_cache.z))
200-element StructArray(::Vector{Float32}, ::Vector{Float32}, ::Vector{Float32}) with eltype SVector{3, Float32}:
 [0.54684204, 0.0305565, 0.19615191]
 [0.12136942, 0.9855579, 0.84889585]
...

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 1, 2025, 2:11pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/25 "2025-10-01T14:11:08Z")

</div>

I’m settling down with this version, based on Zentrik’s suggestion:

```julia-auto
using SIMD: VecRange
using StaticArrays
using BenchmarkTools

function update_dot_exposure!(deltaxy, dot_cache, exposed_i, rj_sq, ::Val{N}) where {N}
    lastN = N * (length(exposed_i) ÷ N)
    lane = VecRange{N}(0)
    @inbounds for i in 1:N:lastN
        if any(exposed_i[lane + i])
            pos_x = dot_cache.x[lane + i] + deltaxy[1]
            pos_y = dot_cache.y[lane + i] + deltaxy[2]
            pos_z = dot_cache.z[lane + i] + deltaxy[3]
            exposed_i[lane + i] &= sum(abs2, (pos_x, pos_y, pos_z)) >= rj_sq
        end
    end
    # Remaining 
    @inbounds for i in lastN+1:length(exposed_i)
        pos_x = dot_cache.x[i] + deltaxy[1]
        pos_y = dot_cache.y[i] + deltaxy[2]
        pos_z = dot_cache.z[i] + deltaxy[3]
        exposed_i[i] &= sum(abs2, (pos_x, pos_y, pos_z)) >= rj_sq
    end
    return exposed_i
end

struct DotCache{T}
    x::Vector{T}
    y::Vector{T}
    z::Vector{T}
end

function data_soa()
    x = rand(SVector{3,Float32}) 
    y = rand(SVector{3,Float32})
    dot_cache = DotCache(rand(Float32,200), rand(Float32,200), rand(Float32,200))
    exposed_i=[rand() > 0.8 for _ in 1:200]
    rj_sq=0.1f0
    N=Val(16)
    return x-y, dot_cache, exposed_i, rj_sq, N
end

```

With the following performance:

```julia-repl
julia> @benchmark update_dot_exposure!($(data_soa())...)
BenchmarkTools.Trial: 10000 samples with 996 evaluations per sample.
 Range (min … max): 22.367 ns … 47.081 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 23.172 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 23.140 ns ± 0.437 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

             ▅▆ ▃█▆ ▂ ▂
  ▇▇▁▁▁▁▁▁▁▃▁███▅▆▄▄▅▆▅▄███▆▆▄▄▅▆▄▅▅▆██▄▆▆▆▅▄▅▅▅▅▇▆▅▇███▇▆▆▆▇ █
  22.4 ns Histogram: log(frequency) by time 24.4 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

Feel free to squeeze it further 🙂

Thanks all for the help!

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 1, 2025, 11:00pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/26 "2025-10-01T23:00:14Z")

</div>

If one consistently uses the fixed-length vector types from SmallCollections.jl (disclaimer: I’m the author), then the easy-peasy broadcasting version becomes noticeably faster:

```julia-auto
julia> @b update_dot_exposure!($(data_soa())...) # lmiq's latest version
28.198 ns

julia> @b update_dot_exposure_small!($(data_soa_small())...) 
18.962 ns

```

> **Code**
>
> ```julia-auto
> using SmallCollections: FixedVector, MutableFixedVector
> 
> struct DotCacheSmall{V}
> x::V
> y::V
> z::V
> end
> 
> function data_soa_small()
> x = rand(FixedVector{3,Float32})
> y = rand(FixedVector{3,Float32})
> M = 200
> dot_cache = DotCacheSmall((FixedVector{M}(rand(Float32, M)) for _ in 1:3)...)
> exposed_i = MutableFixedVector{M}([rand() > 0.8 for _ in 1:M])
> rj_sq = 0.1f0
> return x-y, dot_cache, exposed_i, rj_sq
> end
> 
> sumabs2(x, y, z) = @fastmath abs2(x) + abs2(y) + abs2(z)
> 
> function update_dot_exposure_small!(deltaxy, dot_cache, exposed_i, rj_sq)
> @fastmath exposed_i .&= sumabs2.(deltaxy[1] .+ dot_cache.x, deltaxy[2] .+ dot_cache.y, deltaxy[3] .+ dot_cache.z) .>= rj_sq
> return nothing
> end
> 
> ```

Some comments:

- This requires the current GitHub version `SmallCollections#master`. The published version has a missing `@inline` that spoils performance.
- For some reason this doesn’t work with StaticArrays.jl.
- Now the length of the data vectors is _fixed_. This is of course not what one wants, but one might be able to chop the data into chunks of fixed size, say 128 or 256.

EDIT: Here is a proof of concept with a wrapper `ChunkedVector` (with chunk size 64). This now works for arbitrary data vectors.

> **New code**
>
> ```julia-auto
> const M = 200 # length of data vectors
> const C = 64 # chunk size
> 
> using SmallCollections: FixedVector, MutableFixedVector
> 
> # chunked vectors
> 
> struct ChunkedVector{N,T} <: AbstractVector{T}
> chunks::Vector{MutableFixedVector{N,T}}
> n::Int
> end
> 
> function ChunkedVector{N}(w::AbstractVector{T}) where {N,T}
> m = (length(w)-1) ÷ N + 1
> v = ChunkedVector([zero(MutableFixedVector{N,T}) for _ in 1:m], length(w))
> copyto!(v, w)
> end
> 
> Base.size(v::ChunkedVector) = (v.n,)
> 
> function Base.getindex(v::ChunkedVector{N}, i::Int) where N
> p, q = divrem(i-1, N)
> v.chunks[p+1][q+1]
> end
> 
> function Base.setindex!(v::ChunkedVector{N}, x, i::Int) where N
> p, q = divrem(i-1, N)
> v.chunks[p+1][q+1] = x
> end
> 
> #
> 
> struct DotCacheSmall{V}
> x::V
> y::V
> z::V
> end
> 
> sumabs2(x, y, z) = @fastmath abs2(x) + abs2(y) + abs2(z)
> 
> function update_dot_exposure_chunked!(deltaxy, dot_cache, exposed_i, rj_sq)
> for (x, y, z, e) in zip(dot_cache.x.chunks, dot_cache.y.chunks, dot_cache.z.chunks, exposed_i.chunks)
> @fastmath e .&= sumabs2.(deltaxy[1] .+ x, deltaxy[2] .+ y, deltaxy[3] .+ z) .>= rj_sq
> end
> end
> 
> function data_soa_chunked()
> x = rand(FixedVector{3,Float32})
> y = rand(FixedVector{3,Float32})
> dot_cache = DotCacheSmall((ChunkedVector{C}(rand(Float32, M)) for _ in 1:3)...)
> exposed_i = ChunkedVector{C}([rand() > 0.8 for _ in 1:M])
> rj_sq = 0.1f0
> return x-y, dot_cache, exposed_i, rj_sq
> end
> 
> ```

Performance varies depending on the length. For vectors of length 200 I get

```julia-auto
julia> @b update_dot_exposure!($(data_soa())...)
28.421 ns

julia> @b update_dot_exposure_chunked!($(data_soa_chunked())...)
25.395 ns

```

and for length 300

```julia-auto
julia> @b update_dot_exposure!($(data_soa())...) 
40.457 ns

julia> @b update_dot_exposure_chunked!($(data_soa_chunked())...)
31.273 ns

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 2, 2025, 1:22am UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/27 "2025-10-02T01:22:45Z")

</div>

> [@matthias314](#):
>
> Now the length of the data vectors is _fixed_.

You mean the `M` of the first code? I don’t think I need more flexibility than that (only perhaps because it might stress the compiler if the length varies across executions?)

Ps. Is 200 a small collection? Or we might start having compilation time issues as with StaticArrays?

There was that other package, I think, FixedSizeArrays.jl , can that help?

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 2:33am UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/28 "2025-10-02T02:33:13Z")

</div>

> [@lmiq](#):
>
> You mean the `M` of the first code?

Yes, that’s the length of the data vectors in the first code.

> Is 200 a small collection? Or we might start having compilation time issues as with StaticArrays?

It’s not small in the sense of SmallCollections.jl, so I’m surprised myself that it leads to fast code. But I did notice long compilation times, more than for the chunked version.

> FixedSizeArrays.jl , can that help?

That would be worth trying out!

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 12:52pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/29 "2025-10-02T12:52:32Z")

</div>

> [@lmiq](#):
>
> FixedSizeArrays.jl, can that help?

Turns out that it makes the broadcasting approach much slower.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 2, 2025, 3:25pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/30 "2025-10-02T15:25:09Z")

</div>

Can you open an issue with a minimal reproducer in the repo?

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 2, 2025, 3:54pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/31 "2025-10-02T15:54:08Z")

</div>

But also, much slower…than _ **what** _? What did you swap for what? I’ve been trying to look at this myself, but there are so many different code examples in this thread that I have no idea of what I’m looking at.

If you swapped StaticArrays with FixedSizeArrays, I tried hard to [explain in the docs](https://juliaarrays.github.io/FixedSizeArrays.jl/stable/#Comparison-with-other-array-types) that FixedSizeArrays.jl is _ **not** _ a replacement for StaticArrays.jl, the only property they share is not being able to resize the arrays, for the rest they’re completely different. `FixedSizeArray` is supposed to be very close to `Array`, but not resizable, which is also what the [JuliaCon talk](https://www.youtube.com/watch?v=xWo-ttbSHgw) tried to convey already from the title.

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 4:00pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/32 "2025-10-02T16:00:13Z")

</div>

> <https://github.com/JuliaArrays/FixedSizeArrays.jl/issues/157>
>
> This was observed \[on Discourse\](https://discourse.julialang.org/t/nerd-sniping-…can-you-make-this-faster/132793/29?u=matthias314). @giordano suggested that I report it here.
> 
> In the following code, broadcasting with assignment over \`FixedSizeArray\` is much slower than over \`Vector\`:
> \`\`\`
> using StaticArrays: SVector
> using FixedSizeArrays: FixedSizeArray
> 
> const M = 200
> 
> struct DotCache{V}
> x::V
> y::V
> z::V
> end
> 
> function data\_soa\_vector()
> x = rand(SVector{3,Float32})
> y = rand(SVector{3,Float32})
> dot\_cache = DotCache((rand(Float32, M) for \_ in 1:3)...)
> exposed\_i = \[rand() \> 0.8 for \_ in 1:M \]
> rj\_sq = 0.1f0
> return x-y, dot\_cache, exposed\_i, rj\_sq
> end
> 
> function data\_soa\_fixed()
> # converting from Vector to FixedSizeArray
> deltaxy, dot\_cache, exposed\_i, rj\_sq = data\_soa\_vector()
> dot\_cache\_fixed = DotCache(FixedSizeArray(dot\_cache.x), FixedSizeArray(dot\_cache.y),FixedSizeArray(dot\_cache.z))
> exposed\_i\_fixed = FixedSizeArray(exposed\_i)
> # exposed\_i\_fixed = exposed\_i # with this line instead the benchmarks are identical
> return deltaxy, dot\_cache\_fixed, exposed\_i\_fixed, rj\_sq
> end
> 
> sumabs2(x, y, z) = @fastmath abs2(x) + abs2(y) + abs2(z)
> 
> function update\_dot\_exposure!(deltaxy, dot\_cache, exposed\_i, rj\_sq)
> @fastmath exposed\_i .&= sumabs2.(deltaxy\[1\] .+ dot\_cache.x, deltaxy\[2\] .+ dot\_cache.y, deltaxy\[3\] .+ dot\_cache.z) .\>= rj\_sq
> return nothing
> end
> \`\`\`
> Benchmarks:
> \`\`\`
> julia\> @b update\_dot\_exposure!($(data\_soa\_vector())...)
> 29.512 ns
> 
> julia\> @b update\_dot\_exposure!($(data\_soa\_fixed())...)
> 176.456 ns
> \`\`\`
> As indicated in the code above, making \`exposed\_i\_fixed\` identical to \`exposed\_i\` fixes the problem.
> \`\`\`
> Julia Version 1.11.7
> Commit f2b3dbda30a (2025-09-08 12:10 UTC)
> Build Info:
> Official https://julialang.org/ release
> Platform Info:
> OS: Linux (x86\_64-linux-gnu)
> CPU: 32 × Intel(R) Xeon(R) Platinum 8480+
> WORD\_SIZE: 64
> LLVM: libLLVM-16.0.6 (ORCJIT, sapphirerapids)
> Threads: 1 default, 0 interactive, 1 GC (on 32 virtual cores)
> \`\`\`

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 4:02pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/33 "2025-10-02T16:02:07Z")

</div>

> [@giordano](#):
>
> much slower…than _ **what** _?

I’ve changed the data vectors (length 200) from `Vector` to `FixedSizeArray`. The `SVector` in the code has not been replaced. See the MWE in the GitHub issue I’ve just created.

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 4:51pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/34 "2025-10-02T16:51:57Z")

</div>

I am observing a ~40% slowdown with Julia 1.12.0-rc3 compared to 1.11.7, across all approaches (except for the one using FixedSizeArrays.jl, where apparently some Julia bugs were fixed). For example, @lmiq’s current solution goes from 26.637 ns to 40.826 ns. Do other people also see this?

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 2, 2025, 4:58pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/35 "2025-10-02T16:58:51Z")

</div>

Yes, I confirm the regression:

```julia-auto
julia> @benchmark update_dot_exposure!($(data_soa())...)
BenchmarkTools.Trial: 10000 samples with 994 evaluations per sample.
 Range (min … max): 31.229 ns … 72.302 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 31.284 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 31.424 ns ± 0.680 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▃█▇▁ ▁ ▂▄ ▁
  ████▅▄▄▁▁▁▃▁▁▃▃▄▃▁▁▃▃▁▁▄▃▄▁▃▃▃▁▃▃▇████▇▆▆▅▅▅▄▄▄▅▄▄▃▃▅▄▄▃▁██ █
  31.2 ns Histogram: log(frequency) by time 32.9 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

```

vs [Nerd-sniping: can you make this faster? - #25 by lmiq](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/25)

edit: although in my actual application it is much smaller:

```bash
% time julia +1.12 --project -e "using PDBTools; @time(sasa(atomic_sasa(read_pdb(\"6co8.pdb\"))))"
  5.331173 seconds (55.80 M allocations: 1.993 GiB, 21.69% gc time, 3.27% compilation time)

real	0m5,899s
user	0m6,837s
sys	0m0,339s

#vs

% time julia +1.11 --project -e "using PDBTools; @time(sasa(atomic_sasa(read_pdb(\"6co8.pdb\"))))"
  4.853135 seconds (55.80 M allocations: 1.987 GiB, 17.68% gc time, 0.22% compilation time: 100% of which was recompilation)

real	0m5,400s
user	0m6,273s
sys	0m0,391s

```

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [October 2, 2025, 5:09pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/36 "2025-10-02T17:09:29Z")

</div>

```julia-auto
julia> @benchmark update_dot_exposure!($(data_soa())...)
BenchmarkTools.Trial: 10000 samples with 991 evaluations per sample.
 Range (min … max): 39.168 ns … 130.333 ns ┊ GC (min … max): 0.00% … 0.00%
 Time (median): 44.017 ns ┊ GC (median): 0.00%
 Time (mean ± σ): 45.130 ns ± 6.272 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

  ▁ ▆█▅▂▃▃▁▂ ▁
  █▆▇█████████▆▆▆▅▆▆▇▆▆▇▆▆▆▆▅▅▅▅▅▅▅▅▄▄▃▃▃▃▃▁▄▁▃▁▁▁▃▃▁▁▁▄▁▁▁▃▁█ █
  39.2 ns Histogram: log(frequency) by time 91 ns <

 Memory estimate: 0 bytes, allocs estimate: 0.

julia> versioninfo()
Julia Version 1.12.0-rc3
Commit 7522b240144 (2025-09-26 07:42 UTC)
Build Info:
  Official https://julialang.org release
Platform Info:
  OS: Linux (x86_64-linux-gnu)
  CPU: 14 × Intel(R) Core(TM) Ultra 7 155U
  WORD_SIZE: 64
  LLVM: libLLVM-18.1.7 (ORCJIT, alderlake)
  GC: Built with stock GC
Threads: 14 default, 1 interactive, 14 GC (on 14 virtual cores)

```

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [October 2, 2025, 5:15pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/37 "2025-10-02T17:15:03Z")

</div>

> [@lmiq](#):
>
> in my actual application it is much smaller

Out of curiosity: How does the broadcasting approach compare to SIMD in 1.12 for the whole application? For `update_dot_exposure!` I cannot see any difference anymore.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 2, 2025, 5:26pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/38 "2025-10-02T17:26:54Z")

</div>

It is somewhat slower:

```julia-auto
% time julia +1.12 --project -e "using PDBTools; @time(sasa(atomic_sasa(read_pdb(\"6co8.pdb\"))))"
  6.570915 seconds (57.20 M allocations: 2.143 GiB, 15.10% gc time, 5.72% compilation time)

real	0m8,321s
user	0m9,217s
sys	0m0,380s

% time julia +1.11 --project -e "using PDBTools; @time(sasa(atomic_sasa(read_pdb(\"6co8.pdb\"))))"
  6.444734 seconds (57.46 M allocations: 2.149 GiB, 12.27% gc time, 8.98% compilation time: 2% of which was recompilation)

real	0m7,656s
user	0m8,527s
sys	0m0,394s

```

Or, to be a little bit more clear:

```julia-auto
julia> @b sasa(atomic_sasa(read_pdb("6co8.pdb")))
6.005 s (55785174 allocs: 2.071 GiB, 10.89% gc time, without a warmup)

```

vs

```julia-auto
julia> @b sasa(atomic_sasa(read_pdb("6co8.pdb")))
5.005 s (55785174 allocs: 1.992 GiB, 16.77% gc time, without a warmup)

```

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 2, 2025, 9:16pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793/39 "2025-10-02T21:16:14Z")

</div>

For the record, it [turned out](https://github.com/JuliaArrays/FixedSizeArrays.jl/issues/157) that the slowdown `Vector` vs `FixedSizeVector` was real on Julia v1.11, but that’s gone in Julia v1.13 (and actually v1.12 too, I just tested). There are a bunch of compiler issues that we discovered during the development of the package that have been fixed between 1.12 and 1.13.

[Previous page](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793.md?page=1)
