# Multithreaded code on beefy computer runs just as fast as serial code on M1 Mac

**URL:** <https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210>\
**Category:** Performance\
**Tags:** performance, parallel, differentialequation\
**Created:** [January 26, 2022, 2:30am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210 "2022-01-26T02:30:17Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 2:30am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/1 "2022-01-26T02:30:17Z")

</div>

I have 2 machines:

1. M1 Mac Mini, with horrible thermals and a peak of 3.2 GHz and 8 GB of RAM (absolutely love it though)
2. A beefy linux (Ubuntu 21.10) workstation with 64 GB of DDR5 RAM and a 12th-Gen 4.9 GHz processor with 16 threads. (8 physical cores, efficiency cores disabled)

I have some code, seen [here](https://github.com/oashour/PatternFormation.jl/blob/main/src/PDEDefs.jl), that is used as a function to solve a PDE. This is the parallel version of the code, to get the serial version you obviously just remove `Threads.@spawn` and `@sync`.

Running `example.jl` from that repo and timing the DifferentialEquations.jl `solve` function, we get:

| | Mac | Work Station Serial | Work Station FLoops |
| --- | --- | --- | --- |
| t (s) | ~60 | ~60 | ~60 |

As you can see, the performance is basically the same between the two machines (serial version), and the parallel version runs at the same speed as well after moving to FLoops.

I also tried

```julia
@btime rand(10000,10000)*rand(10000,10000)

```

The M1 takes 11.591 s and the Workstation takes 9.730 s, almost the same. Something is going on.

Here are my questions:

1. (General) does higher clock speed always equal quicker code if everything is OK? Or are there some situations where some bottlenecks won’t be overcome by clock speed?
2. How do I improve the serial performance? Why is a water-cooled 4.9 GHz processor barely keeping up with a 3.2 GHz processor with poor thermals?
3. What’s wrong with my parallel implementation? I tried LoopVectorization.jl with `@tturbo` but couldn’t get it working at all.
4. Is there some good benchmarking code for this?

Thanks in advance, I hope this brings up some interesting discussion and learning opportunities.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [January 26, 2022, 2:34am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/2 "2022-01-26T02:34:11Z")

</div>

Have you tried writing this code with `DifferentialEquations.jl` I’d expect it to be a few orders of magnitude better since it knows fancier time stepping algorithms than you do.

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 2:42am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/3 "2022-01-26T02:42:46Z")

</div>

The time solution is done using DifferentialEquations.jl, see [here](https://github.com/oashour/PatternFormation.jl/blob/main/src/example.jl) for how I do it.

This part is the spatial discretization, which is actually done in this devectorized way for performance. For example, see [here](https://diffeq.sciml.ai/stable/tutorials/faster_ode_example). Writing it this way is 10x quicker than using DiffEqOperators or the matrix form of the Laplacian.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [January 26, 2022, 3:05am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/4 "2022-01-26T03:05:55Z")

</div>

Don’t use `@spawn` like this. Use `Threads.@threads` or `FLoops.@floop`.

For an introduction to multicore parallelism in Julia, see [Data-parallel programming in Julia](https://juliafolds.github.io/data-parallelism/)

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 4:14am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/5 "2022-01-26T04:14:56Z")

</div>

Thank you for the advice. I tried doing it like [this](https://github.com/oashour/PatternFormation.jl/blob/main/src/PDEDefs.jl):

```julia
function GS_Neumann0!(du,u,p,t) # Works only with square grids.
  local f, k, D₁, D₂, dx, dy, M = p
  local N = Int(M)

  @floop for j in 2:N-1, i in 2:N-1
    du[i,j,1] = D₁*(1/dx^2*(u[i-1,j,1] + u[i+1,j,1] - 2u[i,j,1])+ 1/dy^2*(u[i,j+1,1] + u[i,j-1,1] - 2u[i,j,1])) +
                -u[i,j,1]*u[i,j,2]^2 + f*(1-u[i,j,1])
  end
  @floop for j in 2:N-1
    local i = N
    du[i,j,1] = D₁*(1/dx^2*(2u[i-1,j,1] - 2u[i,j,1])+ 1/dy^2*(u[i,j+1,1] + u[i,j-1,1] - 2u[i,j,1])) +
                -u[i,j,1]*u[i,j,2]^2 + f*(1-u[i,j,1])
  end
  # Many more of these

```

and I get the warning/error:

```julia
 Warning: Correctness and/or performance problem detected
│ error =
│ HasBoxedVariableError: Closure ##reducing_function#389#115 (defined in PatternFormation) has 1 boxed variable: N

```

I haven’t quite been able to figure it out for N  
, but fixed it for i and j using `local`. This code runs in ~60 seconds, just like the serial version, so there should be a ton of room for improvement.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [January 26, 2022, 4:20am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/6 "2022-01-26T04:20:50Z")

</div>

You are hitting this case: [https://juliafolds.github.io/FLoops.jl/dev/explanation/faq/#uncertain-values](https://juliafolds.github.io/FLoops.jl/dev/explanation/faq/#uncertain-values)

This is the problem:

[https://github.com/oashour/PatternFormation.jl/blob/1768c25b1f465b3980f9f1c7baaf865b198cea23/src/PDEDefs.jl#L2-L4](https://github.com/oashour/PatternFormation.jl/blob/1768c25b1f465b3980f9f1c7baaf865b198cea23/src/PDEDefs.jl#L2-L4)

Either use `let N = N ... ` as mentioned in the FAQ or use

```julia
  f, k, D₁, D₂, dx, dy, Ntmp = p
  N = Int(Ntmp)

```

i.e., introduce `Ntmp` to avoid setting `N` twice

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 4:31am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/7 "2022-01-26T04:31:29Z")

</div>

Thank you, `let` fixed the error. The code still runs in ~60 seconds same as the sequential so I have to work on figuring that out (assuming I can figure out why the Workstation and the Mac run at the same wall time).

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [January 26, 2022, 4:46am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/8 "2022-01-26T04:46:12Z")

</div>

You can use `--threads=` with different numbers and see how the run-time change in the workstation. Also, monitor `top`/`htop`/etc. output see if CPUs are active as you’d expect. If the run-time does not change with varying `--threads=`, it’s likely that the bottleneck is elsewhere.

Bonus: you can also make the executor explicit as in

```julia
function GS_Periodic!(du,u,p,t; ex = nothing)
  ...

  @floop ex for ...

```

and pass `ex = ThreadedEx(basesize = ...)` to change the number of “effective threads” [https://juliafolds.github.io/data-parallelism/howto/faq/#can\_i\_change\_the\_number\_of\_execution\_threads\_without\_restarting\_julia](https://juliafolds.github.io/data-parallelism/howto/faq/#can_i_change_the_number_of_execution_threads_without_restarting_julia)

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 5:10am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/9 "2022-01-26T05:10:52Z")

</div>

I tried doing it like this:

```julia
N = 256 # Grid is 256x256, each un-nested for loop goes over N-2 elements
N_threads = 16
bs = (N -2)÷ N_threads
ex = ThreadedEx(basesize = bs)

```

And for N = 4,8,16, they all take 60 seconds. I checked the processor utilization with `top` and it seems that it mostly hovers at 100% and sometimes jumps to 200%, but that’s it. I assume this means it is not properly using the 16 threads I have access to. /sigh

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [January 26, 2022, 5:18am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/10 "2022-01-26T05:18:15Z")

</div>

You need to launch julia with `julia --threads=16`

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 5:29am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/11 "2022-01-26T05:29:04Z")

</div>

Thank you, I have already done this prior and changed it in the VSCode settings for the remote machine (the Workstation). Running `Threads.nthreads()` returns 16.

As suggested by @tkf, I ran `htop` and monitored Julia spawning the threads, it’s just that most threads use no computing power whatsoever. Weird.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [January 26, 2022, 5:48am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/12 "2022-01-26T05:48:02Z")

</div>

I just skimmed the code bit more. You have a lot of `for` loops but they don’t depend on each other and the iteration space are all the same. So, IIUC, it looks like you can rewrite them to use one `for j in 2:N-1, i in 2:N-1` loop and one `for j in 2:N-1` loop? Also, if `N = 256` is a typical size, I’d use `@floop` only for the nested loop.

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 7:53am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/13 "2022-01-26T07:53:53Z")

</div>

I changed the code, partially implementing your suggestions. It did not make an appreciable difference in the results. Then I changed the linear solver of DifferentialEquations.jl from `linsolve=KLUFactorization` to whatever the default is. The results are in the table below:

| | WS 1t | WS 8t | M1 1t | M1 8t |
| --- | --- | --- | --- | --- |
| t (s) | 35.58 | 36.24 | 70.34 | 73.6 |

where WS is the workstation and M1 is the Mac. 1t is 1 thread (sequential execution). As you can see, the WS is now twice as fast! But the parallelization makes basically no difference.

So now I have to figure out why the parallelization is not working as it should.

@ChrisRackauckas sorry for the direct ping, but would you happen to know why KLU factorization causes this issue? (tl;dr KLU factorization runs at the same speed on two machines, one fast one slow. Default algorithm is twice as fast on fast machine).

---

<div class="post-metadata">

**Author:** ![trahflow](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/trahflow/32/30585_2.png) [@trahflow](https://discourse.julialang.org/u/trahflow)\
**Post date:** [January 26, 2022, 9:11am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/14 "2022-01-26T09:11:11Z")

</div>

I know this does not directly address your question, and maybe you know the resource already, but if not this lecture by Chris Rackauckas is really an awesome resource for improving ones understanding of exactly these kinds of problems: [https://www.youtube.com/playlist?list=PLCAl7tjCwWyGjdzOOnlbGnVNZk0kB8VSa](https://www.youtube.com/playlist?list=PLCAl7tjCwWyGjdzOOnlbGnVNZk0kB8VSa)

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 9:40am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/15 "2022-01-26T09:40:54Z")

</div>

Thank you, I wasn’t aware of this actually. Looks fantastic!

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [January 26, 2022, 9:46am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/16 "2022-01-26T09:46:08Z")

</div>

> [@ash](#):
>
> @ChrisRackauckas sorry for the direct ping, but would you happen to know why KLU factorization causes this issue? (tl;dr KLU factorization runs at the same speed on two machines, one fast one slow. Default algorithm is twice as fast on fast machine).

Did you profile? How much time is spent in the KLU? IIRC, KLU is a fairly serial factorization algorithm for sparsity patterns with low amounts of structure. The default (UMFPACK) can exploit more structure and threading, but requires some repeated structures in the sparsity pattern. Which one will be better depends on the problem and the compute hardware.

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 10:11am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/17 "2022-01-26T10:11:33Z")

</div>

I ran a little experiment to figure out the effects of multithreading on both systems. I extended the simulation time by 100 fold and ran 2 simulations, one with 1 thread and one with 16 threads on the workstation or 8 on the Mac. I also used `BLAS.set_num_threads(N_threads)`. I additionally started using MKL, which didn’t make much of a difference, unfortunately.

The run times are in the table below. As you can see, they are all more or less the same, so we’re back to square zero:

1. The multithreaded code is as fast as the sequential code on the WS
2. Parallelism cripples the Mac for some reason, but I am more worried about debugging the workstation’s performance.
3. The tiny M1 Mac Mini runs the code as fast as the workstation, even when using 16 threads on the workstation and 1 on the Mac.

| | WS 1t | WS 16t | M1 1t | M1 8t |
| --- | --- | --- | --- | --- |
| t (s) | 327 | 338 | 333 | 546 |

I was monitoring `htop` the whole time, and for the single threaded simulation, one thread stayed at 100% the whole time. On the other hand, in the multithreaded calculation, all 16 threads jumped around between 0 and 50%, with an average overall CPU usage of around ~400-500%.

I think it makes sense the multithreaded and sequential versions have the same run time based on this data, accounting for the parallelism overhead.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [January 26, 2022, 10:50am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/18 "2022-01-26T10:50:37Z")

</div>

Could you post an annotated profile?

---

<div class="post-metadata">

**Author:** ![ash](https://avatars.discourse-cdn.com/v4/letter/a/f19dbf/32.png) [@ash](https://discourse.julialang.org/u/ash)\
**Post date:** [January 26, 2022, 11:20am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/19 "2022-01-26T11:20:32Z")

</div>

Embarrassingly, I don’t know much about profiling and don’t know how to annotate one.

However, I profiled some code on both systems and tried UMF vs KLU. There is a weird bug where `@profview` from VSCode does not capture the true time it takes to run the function vs `@time`. Thus, the M1 results below aren’t reliable until I figure out how to profile it more properly.

| | WS | M1 |
| --- | --- | --- |
| Total | 59 | 39 (55) |
| `KLU.LibKLU.klu_l_factor` | 46.1 | 33.1 |
| `KLU.LibKLU.klu_l_solve` | 8.7 | 0.8 |

| | WS | M1 |
| --- | --- | --- |
| Total` | 32 | 14 (33) |
| `SuiteSparse.UMFPACK.solve!` | 20.9 | 1.8 |
| `SuiteSparse.UMFPACK.#umfpack_numeric!#13` | 4.9 | 2.6 |
| `SuiteSparse.UMFPACK.umfpack_symbolic`! | 2.22 | 2.2 |

The number between parentheses for the M1 is the actual run time if I use `@time`. So basically both systems perform identically now with UMF being almost twice as fast.

Unfortunately, I don’t think this data is very useful due to the M1 bug, but at least it can tell us more about the WS.

EDIT: all tests were done with 1 thread and OpenBLAS.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [January 26, 2022, 11:30am UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210/20 "2022-01-26T11:30:58Z")

</div>

[![](https://global.discourse-cdn.com/julialang/original/3X/0/9/0947182566549a263febea8b3de0890dbd59f0e1.jpeg "Code Profiling and Optimization (in Julia)") ](https://www.youtube.com/watch?v=h-xVBD2Pk9o)

shows how to profile. Profile a small solve. Don’t optimize what the profile doesn’t show.

[Next page](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210.md?page=2)
