# Parallel computation of multiplication of large matrices

**URL:** <https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898>\
**Category:** Performance\
**Created:** [April 8, 2019, 12:55am UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898 "2019-04-08T00:55:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [April 8, 2019, 12:55am UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/1 "2019-04-08T00:55:29Z")

</div>

A common task in many programs is to multiply two large matrices, namely, A = B \*C. Is there a way to do it in a way that utlizes all the processors ? This is a task that is theoretically, very easy to parallelize over any number of threads.

I would appreciate any help on this.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [April 8, 2019, 2:38am UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/2 "2019-04-08T02:38:45Z")

</div>

Over threads? `A = B*C`. Try it!

```julia
A = rand(3000,3000); B = rand(3000,3000)
A*B

```

Open up your resource manager and see the usage. Notice that this is using OpenBLAS (or MKL) under the hood, which is already multithreading. The number of threads matches the number of physical cores by default. You can override this with

```julia
using LinearAlgebra
BLAS.set_num_threads(i)

```

For distributed computations over multiple computers, you’ll need to use multiprocessing. For that, the simplest way to get a usable solution is to make `A` and `B` be MPIArrays from MPIArrays.jl:

> **[GitHub - barche/MPIArrays.jl: Distributed arrays based on MPI onesided...](https://github.com/barche/MPIArrays.jl)**
>
> Distributed arrays based on MPI onesided communication - GitHub - barche/MPIArrays.jl: Distributed arrays based on MPI onesided communication

Then `A*B` between two MPI arrays will be parallelized via MPI. And to complete this response, you can use GPUArrays.jl/CuArrays.jl to do

> **[GitHub - JuliaGPU/CuArrays.jl: A Curious Cumulation of CUDA Cuisine](https://github.com/JuliaGPU/CuArrays.jl)**
>
> A Curious Cumulation of CUDA Cuisine. Contribute to JuliaGPU/CuArrays.jl development by creating an account on GitHub.

```julia
_A = cu(A); _B = cu(B)
_A*_B

```

to parallelize it on the GPU.

---

<div class="post-metadata">

**Author:** ![traktofon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/traktofon/32/591_2.png) [@traktofon](https://discourse.julialang.org/u/traktofon)\
**Post date:** [April 8, 2019, 11:18am UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/3 "2019-04-08T11:18:03Z")

</div>

> [@ChrisRackauckas](#):
>
> The number of threads matches the number of physical cores by default.

I thought OpenBLAS is using max 16 threads? (at least for the precompiled Julia versions)

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [April 8, 2019, 11:26am UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/4 "2019-04-08T11:26:41Z")

</div>

> [@traktofon](#):
>
> I thought OpenBLAS is using max 16 threads? (at least for the precompiled Julia versions)

Yes, there is the caveat that if you have more than 16 physical cores the default will be capped to 16, so you’ll need to build Julia from source. I assume that’s not a very standard concern.

---

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [April 8, 2019, 3:25pm UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/5 "2019-04-08T15:25:48Z")

</div>

Thanks, that was very very enlightening

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 8, 2019, 4:07pm UTC](https://discourse.julialang.org/t/parallel-computation-of-multiplication-of-large-matrices/22898/6 "2019-04-08T16:07:35Z")

</div>

Ooooh… I might have access to an ARM machine with 50 cores. I might not.  
Might be fun to see this behaviour.

Actually I did try to build Julia from surce on that machine and it might have fork bombed ….  
Not sure if it was Julia or the CEPH I was trying to compile at the same time.
