# Non-intuitive perf diff between \`matrix \* vector\`, \`matrix' \* vector\` and \`copy(matrix') \* vector\`

**URL:** <https://discourse.julialang.org/t/non-intuitive-perf-diff-between-matrix-vector-matrix-vector-and-copy-matrix-vector/29207>\
**Category:** Performance\
**Tags:** blas\
**Created:** [September 26, 2019, 5:43pm UTC](https://discourse.julialang.org/t/non-intuitive-perf-diff-between-matrix-vector-matrix-vector-and-copy-matrix-vector/29207 "2019-09-26T17:43:37Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jeremejevs](https://avatars.discourse-cdn.com/v4/letter/j/ea666f/32.png) [@jeremejevs](https://discourse.julialang.org/u/jeremejevs)\
**Post date:** [September 26, 2019, 5:43pm UTC](https://discourse.julialang.org/t/non-intuitive-perf-diff-between-matrix-vector-matrix-vector-and-copy-matrix-vector/29207/1 "2019-09-26T17:43:37Z")

</div>

I’m using Julia 1.2. This is my test:

```julia
a = rand(1000, 1000)
b = a'
c = copy(b)

@btime a * x setup=(x=rand(1000)) # 114.757 μs
@btime b * x setup=(x=rand(1000)) # 94.179 μs
@btime c * x setup=(x=rand(1000)) # 110.325 μs

```

I was expecting **a** and **c** to be at very least not slower.

After inspecting **stdlib/LinearAlgebra/src/matmul.jl** , it turns out that Julia passes **b.parent** (i.e. **a** ) to **BLAS.gemv** , not **b** , and instead switches LAPACK’s **dgemv\_** into a different and apparently faster mode.

Am I correct in assuming that the speedup comes from the fact that the memory is aligned in a more favorable way for whatever **dgemv\_** does, when it’s in a **trans = T** mode? If so, then I’m guessing this isn’t actionable, besides possibly mentioning the gotcha in the docs somehow. If my assumption is wrong though, is there something to be done about this?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 26, 2019, 5:52pm UTC](https://discourse.julialang.org/t/non-intuitive-perf-diff-between-matrix-vector-matrix-vector-and-copy-matrix-vector/29207/2 "2019-09-26T17:52:11Z")

</div>

> [@jeremejevs](#):
>
> Am I correct in assuming that the speedup comes from the fact that the memory is aligned in a more favorable way for whatever **dgemv\_** does, when it’s in a **trans = T** mode?

Close. It does have to do with memory, but it’s about [locality](https://en.wikipedia.org/wiki/Locality_of_reference), not alignment. The basic thing to understand is that it is more efficient to access _consecutive_ (or at least nearby) data from memory than data that is separated, due to the existence of [cache lines](https://en.wikipedia.org/wiki/CPU_cache#Cache_entries). (Consecutive access also has some advantages in utilizing [SIMD](https://en.wikipedia.org/wiki/SIMD) instructions.)

Julia stores matrices in [column-major order](https://en.wikipedia.org/wiki/Row-_and_column-major_order), so that the columns are contiguous in memory. When you multiply a _transposed_ matrix (that has not been copied) by a vector, therefore, it can compute it as the dot product of the contiguous column (= transposed row) with the contiguous vector, which has good **spatial locality** and therefore utilizes cache lines efficiently.

For multiplying a _non_-transposed matrix by a vector, in contrast, you are taking the dot products of non-contiguous _rows_ of the matrix with the vector, and it is harder to efficiently utilize cache lines. To improve spatial locality in this case, an optimized BLAS like OpenBLAS actually computes the dot products of several rows at a time (a “block”) with the vector, I believe — that’s why it’s only 10% slower and not much worse. (In fact, even the transposed case may do some blocking to keep the vector in cache.)

---

<div class="post-metadata">

**Author:** ![jeremejevs](https://avatars.discourse-cdn.com/v4/letter/j/ea666f/32.png) [@jeremejevs](https://discourse.julialang.org/u/jeremejevs)\
**Post date:** [September 27, 2019, 7:09am UTC](https://discourse.julialang.org/t/non-intuitive-perf-diff-between-matrix-vector-matrix-vector-and-copy-matrix-vector/29207/3 "2019-09-27T07:09:25Z")

</div>

Couldn’t have wished for a better reply, thank you!
