# Improving Base.mapreduce by 20% on vectors

**URL:** <https://discourse.julialang.org/t/improving-base-mapreduce-by-20-on-vectors/134478>\
**Category:** Performance\
**Tags:** performance, pluto, mapreduce\
**Created:** [December 10, 2025, 4:43pm UTC](https://discourse.julialang.org/t/improving-base-mapreduce-by-20-on-vectors/134478 "2025-12-10T16:43:34Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [December 10, 2025, 4:43pm UTC](https://discourse.julialang.org/t/improving-base-mapreduce-by-20-on-vectors/134478/1 "2025-12-10T16:43:34Z")

</div>

Hello everyone,

I’ve been exploring reduction operations on GPUs recently, and this led me to investigate CPU performance as well. I discovered that we can improve `mapreduce` in Base while maintaining full precision for floating-point operations.

My analysis suggests there may be an alignment issue with SIMD in the current Base implementation, which leads to suboptimal performance for common computational patterns. Additionally, I’ve implemented specialized functions for small arrays (size \< 32) that provide significant speedups.

https ://epilliat.github.io/coding/notebooks/mapreduce\_cpu\_perf.html

( **remove the space** after https), I cannot put a link to the html file nor upload it.

I’ve documented the entire analysis and implementation in a Pluto notebook:

More details in this notebook will be added later.

The key results:

- ~20% performance improvement over Base for most array sizes
- Maintained numerical precision for floating-point types
- Better performance across multiple numeric types including Float32

I’d appreciate any feedback on the approach and results!

---

<div class="post-metadata">

**Author:** ![adienes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adienes/32/37459_2.png) [@adienes](https://discourse.julialang.org/u/adienes)\
**Post date:** [December 10, 2025, 4:52pm UTC](https://discourse.julialang.org/t/improving-base-mapreduce-by-20-on-vectors/134478/2 "2025-12-10T16:52:35Z")

</div>

looks like a fun experiment!

you may be interested in this thread: [Performance challenge: can you write a faster sum?](https://discourse.julialang.org/t/performance-challenge-can-you-write-a-faster-sum/130456)

and there are lots of benchmarks in the associated PR [WIP: The great pairwise reduction refactor by mbauman · Pull Request #58418 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/pull/58418)

the most challenging part here is getting (as) uniform (as possible) speedups across all array types, shapes, sizes, element types, computer architectures, etc.

for example, a change to `mapreduce` that makes it 20% faster on `Array{Float64}` might accidentally cause 10x regressions on a `ReshapedArray{BigInt, 2, SubArray{...}}` (not particularly that type, just made something up for dramatic effect)

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [December 12, 2025, 1:32pm UTC](https://discourse.julialang.org/t/improving-base-mapreduce-by-20-on-vectors/134478/3 "2025-12-12T13:32:17Z")

</div>

Thank you very much for the references !

I’m getting into the post you mention. In any case, this is a very interesting problem I guess, both for optimizing Julia and for learning new stuff about Julia.

As you say, using @simd looks to be a performance killer for views and more complex types. However, I cannot see how to do better (at least on my computer) than @simd for small concrete types and coalesced memory – see my reply on the performance challenge post.
