# Pmap and svdvals from documentation speed

**URL:** <https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898>\
**Category:** Performance\
**Tags:** question, distributed\
**Created:** [September 8, 2021, 9:24pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898 "2021-09-08T21:24:00Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pulpo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pulpo/32/8502_2.png) [@Pulpo](https://discourse.julialang.org/u/Pulpo)\
**Post date:** [September 8, 2021, 9:24pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/1 "2021-09-08T21:24:00Z")

</div>

I am wondering why the following example of pmap isn’t getting any speedup:

> using Distributed  
> using BenchmarkTools  
> @everywhere using LinearAlgebra
> 
> addprocs(length(Sys.cpu\_info()));  
> M = Matrix{Float64}[rand(1000,1000) for i = 1:10];  
> @btime pmap(svdvals, M) ; # \_\_1.931 s  
> @btime map(svdvals,M) ; # \_\_\_\_764.653 ms

This is an example taken from the Distributed computing [manual](https://docs.julialang.org/en/v1/manual/distributed-computing/). This behavior doesn’t only happen for the svdvals, but also what I actually need it for, backslash. Also, even if I were to compute the svdvals of each M[i] in a forloop, the pmap still seems to be outperformed.

Increasing the size of M doesn’t really seem to improve things. So I am wondering why this is happening and if there is a way to actually get speedup in a situation where I would need to calculate the svd, or backslash, of a matrix of the type M (or an Array{Float64,(NN,N,N)} along NN)?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [September 8, 2021, 9:39pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/2 "2021-09-08T21:39:24Z")

</div>

Most linear algebra stuff in Julia is already multithreaded.

---

<div class="post-metadata">

**Author:** ![Pulpo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pulpo/32/8502_2.png) [@Pulpo](https://discourse.julialang.org/u/Pulpo)\
**Post date:** [September 8, 2021, 10:23pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/3 "2021-09-08T22:23:01Z")

</div>

I understand that, but shouldn’t I see some sort of speedup if instead of

> M = Matrix{Float64}[rand(1000,1000) for i = 1:10];

I increase the size to

> M = Matrix{Float64}[rand(1000,1000) for i = 1:1000];

which doesn’t seem to be the case. I am in a situation where I have a bunch of large linear systems to solve. To my understanding pmap should be able to do this, but if this isn’t the case then what would be the best way of going about this? Using the MPI.jl?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [September 8, 2021, 10:32pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/4 "2021-09-08T22:32:40Z")

</div>

Are you trying this on a multi-computer cluster? If not, I wouldn’t expect it to be faster.

---

<div class="post-metadata">

**Author:** ![Pulpo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pulpo/32/8502_2.png) [@Pulpo](https://discourse.julialang.org/u/Pulpo)\
**Post date:** [September 8, 2021, 10:48pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/5 "2021-09-08T22:48:40Z")

</div>

I have sequential code on the a cluster, but I have not implemented the parallelization yet as I am trying to figure out a (intelligent) way of doing it. I just have been experimenting on my laptop. I should mention that part of the job of each process should also be in building the matrices, they are not random as in the example above.

But you expect that on a cluster I should be able to see some speedup?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [September 8, 2021, 11:00pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/6 "2021-09-08T23:00:12Z")

</div>

For testing on your laptop, set BLAS threads to 1 using `BLAS.set_num_threads(1)`. That said, if you are trying to optimize performance for a cluster, testing on a laptop will not give useful results.

---

<div class="post-metadata">

**Author:** ![Pulpo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pulpo/32/8502_2.png) [@Pulpo](https://discourse.julialang.org/u/Pulpo)\
**Post date:** [September 8, 2021, 11:09pm UTC](https://discourse.julialang.org/t/pmap-and-svdvals-from-documentation-speed/67898/7 "2021-09-08T23:09:46Z")

</div>

I already tried using `BLAS.set_num_threads(1)` , but the speeds are still comparable, which is another reason why I am a little confused by this example from the manual. But seems like I should switch to testing on the cluster according to what you are saying.
