# Why is sum slower than multiplying by vector of ones?

**URL:** <https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347>\
**Category:** Performance\
**Created:** [September 28, 2023, 11:51am UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347 "2023-09-28T11:51:02Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jad\_Zeitouni](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jad_zeitouni/32/50678_2.png) [@Jad\_Zeitouni](https://discourse.julialang.org/u/Jad_Zeitouni)\
**Post date:** [September 28, 2023, 11:51am UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347/1 "2023-09-28T11:51:02Z")

</div>

So when I try to sum a 10000 x 100 matrix `x` on its second dimension, using `sum(x, dims=2)`, it takes about 200 microseconds. If I instead do `x * ones(100)`, this takes only 40 microseconds. What gives?

Here’s the full code

```julia
using BenchmarkTools
x = rand(10000, 100);
@benchmark sum($x, dims=2) # mean ≈ 200 μs
@benchmark $x * ones(100) # mean ≈ 40 μs

```

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [September 28, 2023, 12:52pm UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347/3 "2023-09-28T12:52:52Z")

</div>

it’s adding in an order that takes better advantage of the cache. we should probably make sum use this sort of ordering in base.

---

<div class="post-metadata">

**Author:** ![Jad\_Zeitouni](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jad_zeitouni/32/50678_2.png) [@Jad\_Zeitouni](https://discourse.julialang.org/u/Jad_Zeitouni)\
**Post date:** [September 28, 2023, 1:09pm UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347/4 "2023-09-28T13:09:10Z")

</div>

Would you mind elaborating? As someone who knows very little of lower level programming I’m eager to learn.

Would it even be possible to implement this faster ordering using for loops in Julia? A simple for loop implementation does a bit better than `sum` but not much.

```julia
function f(x)
    m, n = size(x)
    res = zeros(m)
    @simd for j in 1:n
        @simd for i in 1:m
            @fastmath @inbound res[i] += x[i, j]
            end
        end
    end
    return res
end

```

---

<div class="post-metadata">

**Author:** ![HanD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hand/32/213908_2.png) [@HanD](https://discourse.julialang.org/u/HanD)\
**Post date:** [September 28, 2023, 1:42pm UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347/5 "2023-09-28T13:42:47Z")

</div>

I’m guessing this is because matrix operations are handled by BLAS, which uses multiple threads for parallel computations by default. Whereas `sum` is implemented in Julia, and runs on a single thread by default.

Watching CPU usage during the benchmarks seems to confirm this hypothesis. If you set the envvar `OPENBLAS_NUM_THREADS=1` before running the benchmarks, the time difference becomes significantly smaller (albeit it’s still there).

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [September 28, 2023, 2:05pm UTC](https://discourse.julialang.org/t/why-is-sum-slower-than-multiplying-by-vector-of-ones/104347/6 "2023-09-28T14:05:04Z")

</div>

Funny, I was just writing some code and considered whether I should use `sum` or the more math-looking operation and I chose `sum`. Looks like I chose slow!
