# Outperformed by Matlab

**URL:** <https://discourse.julialang.org/t/outperformed-by-matlab/71603>\
**Category:** Performance\
**Tags:** matlab, multithreading, linearalgebra, tullio, loopvectorization\
**Created:** [November 16, 2021, 5:58pm UTC](https://discourse.julialang.org/t/outperformed-by-matlab/71603 "2021-11-16T17:58:00Z")\
**Posts on this page:** 1\
**Showing post:** 48

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [November 19, 2021, 6:37pm UTC](https://discourse.julialang.org/t/outperformed-by-matlab/71603/48 "2021-11-19T18:37:05Z")

</div>

> [@Germinator](#):
>
> Note that `batchSize = 10000;` means that in the current example, you would first form the whole X matrix and then compute X’X and X’y in two separate steps. In practice, I have been considering `batchSize = 10000;` when N≈10^7.

So your requirements are N around 10^7, but R and M stay in their order of magnitude?

> [@Germinator](#):
>
> As in the latest reply to @goerch, this is probably very naive, so my hope was that smart threading of @tturbo and @tullio could do the trick.

LinearAlgebra and BLAS underneath already are multithreaded.

Just tried my latest and greatest for `N = 1000000`:

```julia
@btime normal1!($C,$Cx,$x,$A,$B)

for blocksize in [10, 100, 1000, 10000, 100000]
    @btime normal2!($C,$Cx,$x,$A,$B,$blocksize)
end

```

with results

```julia
  48.886 s (1 allocation: 6.38 KiB)
  10.600 s (7 allocations: 130.25 KiB)
  7.888 s (8 allocations: 1.27 MiB)
  8.257 s (9 allocations: 12.67 MiB)
  8.619 s (10 allocations: 126.72 MiB)
  8.492 s (10 allocations: 1.24 GiB)

```

My 6 cores are running hot. Most time is spent in `gemm`. How fast do you think we can get this?

Edit: obvious improvement: implementing `C += temp * temp'` with rank-1 or rank-k updates via  
`BLAS.syr!` and `BLAS.syrk!`, here the results:

```julia
  64.247 s (1 allocation: 6.38 KiB)
  6.456 s (4 allocations: 125.09 KiB)
  4.752 s (4 allocations: 1.22 MiB)
  5.423 s (4 allocations: 12.21 MiB)
  6.103 s (4 allocations: 122.07 MiB)
  5.422 s (4 allocations: 1.19 GiB)

```

Another edit after [discussions](https://discourse.julialang.org/t/mul-dispatch-to-blas-incomplete/71831) using symmetric matrices for `C`:

```julia
  65.309 s (1 allocation: 6.38 KiB)
  6.326 s (4 allocations: 125.09 KiB)
  4.815 s (4 allocations: 1.22 MiB)
  4.616 s (4 allocations: 12.21 MiB)
  5.026 s (4 allocations: 122.07 MiB)
  4.971 s (4 allocations: 1.19 GiB)

```

---

_[View the full topic](https://discourse.julialang.org/t/outperformed-by-matlab/71603)._
