# Strange multithreaded behaviour on ARM Ampere A1 (Oracle Cloud)

**URL:** <https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490>\
**Category:** Julia at Scale\
**Tags:** multithreading, aarch64, arm\
**Created:** [April 14, 2022, 2:01pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490 "2022-04-14T14:01:11Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dandan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dandan/32/8935_2.png) [@Dandan](https://discourse.julialang.org/u/Dandan)\
**Post date:** [April 14, 2022, 2:01pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/1 "2022-04-14T14:01:11Z")

</div>

I’ve tried to run the following simple benchmark in Oracle Cloud VM on an instance with 4 vCPU of ARM Ampere A1 and 24GB ram:

```julia
using LinearAlgebra

function foo()
	Threads.@threads for i in 1:10000
		H = rand(100, 100) + im * rand(100, 100)
		eigen(Hermitian(H))
	end
end

@time foo()

```

When I use

```julia
julia -t 4

```

the running time is around 3 times longer in comparison to the single threaded test.

In both cases Julia correctly sees the value of Threads.nthreads().

Also, in both cases the CPUs utilization is maximal (400%).

I wonder what is going on?

Julia version 1.7.2 (aarch64)

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 14, 2022, 2:34pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/2 "2022-04-14T14:34:24Z")

</div>

> [@Dandan](#):
>
> 4 vCPU of ARM Ampere A1

If you run `lscpu` at the shell, is there more than one NUMA node?

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 14, 2022, 2:36pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/3 "2022-04-14T14:36:17Z")

</div>

> [@Dandan](#):
>
> `eigen`

You might also want to look at `BLAS.set_num_threads` to set the number of threads that OpenBLAS is trying to uses:

> [@Julia Threads vs BLAS threads](https://discourse.julialang.org/t/julia-threads-vs-blas-threads/8914):
>
> I am experimenting with Julia’s (experimental) multithreading feature recently and like the results so far. One of the problems I need to deal with, involves the multiplication of several pairs (say in the order 5 to 50) of matrices, whose size is average (say linear size in the order 10 - 1000). For that problem, there can be a competition between either using Julia threads (to loop over the different pairs) versus using multithreaded matrix multiplication provided by BLAS, and it will depend …

---

<div class="post-metadata">

**Author:** ![Dandan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dandan/32/8935_2.png) [@Dandan](https://discourse.julialang.org/u/Dandan)\
**Post date:** [April 14, 2022, 2:57pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/4 "2022-04-14T14:57:08Z")

</div>

It’s only 1 NUMA node

```julia
$ lscpu
Architecture: aarch64
CPU op-mode(s): 32-bit, 64-bit
Byte Order: Little Endian
CPU(s): 4
On-line CPU(s) list: 0-3
Thread(s) per core: 1
Core(s) per socket: 4
Socket(s): 1
NUMA node(s): 1
Vendor ID: ARM
Model: 1
Model name: Neoverse-N1
Stepping: r3p1
BogoMIPS: 50.00
NUMA node0 CPU(s): 0-3
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; __user pointer sanitization
Vulnerability Spectre v2: Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Not affected
Flags: fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm lrcpc dc
                                 pop asimddp ssbs

```

`BLAS.get_num_threads()` returns 4 in both cases

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 14, 2022, 2:59pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/5 "2022-04-14T14:59:52Z")

</div>

> [@Dandan](#):
>
> `BLAS.get_num_threads()` returns 4 in both cases

Try setting it to 1 via `BLAS.set_num_threads(1)`

---

<div class="post-metadata">

**Author:** ![Dandan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dandan/32/8935_2.png) [@Dandan](https://discourse.julialang.org/u/Dandan)\
**Post date:** [April 14, 2022, 3:26pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/6 "2022-04-14T15:26:39Z")

</div>

Right, with `BLAS.set_num_threads(1)` and 4 threads the code runs 3x times faster than single threaded.

Though, the single threaded case is still strange.  
With `BLAS.set_num_threads(1)` the CPU utilization becomes 100% in this case, but the running time remains the same!

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 14, 2022, 3:31pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/7 "2022-04-14T15:31:49Z")

</div>

> [@Dandan](#):
>
> With `BLAS.set_num_threads(1)` the CPU utilization becomes 100% in this case, but the running time remains the same!

How are you measuring CPU utilization?

---

<div class="post-metadata">

**Author:** ![Dandan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dandan/32/8935_2.png) [@Dandan](https://discourse.julialang.org/u/Dandan)\
**Post date:** [April 14, 2022, 3:33pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/8 "2022-04-14T15:33:58Z")

</div>

It’s what `top` command shows

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 14, 2022, 3:52pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/9 "2022-04-14T15:52:44Z")

</div>

`htop` may give you a better idea on what is happening on a per processor basis.

![image](https://global.discourse-cdn.com/julialang/original/3X/8/6/86522b0aad8556965e7e65cbe4e62a09754b4f43.png)

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [April 14, 2022, 3:57pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/10 "2022-04-14T15:57:13Z")

</div>

> Neoverse-N1

Julia 1.7 - has [OPENBLAS\_VER := 0.3.13](https://github.com/JuliaLang/julia/blob/release-1.7/deps/Versions.make#L84=)

And check the OpenBLAS [changelog](https://github.com/xianyi/OpenBLAS/blob/develop/Changelog.txt) ( search “neoverse” )

- _“3.16 : fixed missing restore of a register … **Neoverse N1** that could cause spurious failures in e.g. DGEEV”_
- _“3.14 : Fixed the THUNDERX2T99 and **NEOVERSEN1** DNRM2/ZNRM2 kernels for inputs with Inf”_

suggestions:

- check Julia [v1.8.0-beta3 (March 29, 2022)](https://julialang.org/downloads/#upcoming_release)
  - it has [“OPENBLAS\_VER := 0.3.17”](https://github.com/JuliaLang/julia/blob/release-1.8/deps/Versions.make#L82=)

- if it is not enough - recompile Julia 1.8 source with ` -march=armv8.2-a -mtune=neoverse-n1`

---

<div class="post-metadata">

**Author:** ![Dandan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dandan/32/8935_2.png) [@Dandan](https://discourse.julialang.org/u/Dandan)\
**Post date:** [April 14, 2022, 6:56pm UTC](https://discourse.julialang.org/t/strange-multithreaded-behaviour-on-arm-ampere-a1-oracle-cloud/79490/11 "2022-04-14T18:56:17Z")

</div>

I’ve tried beta version and recompilation, but the results are the same.

It’s what `htop` shows when the code runs with the default number of BLAS threads (4).

 ![Screenshot 2022-04-14 at 21.48.11](https://global.discourse-cdn.com/julialang/original/3X/7/0/7054c87850656d7703fe5192cbca9d134e204502.png)

 ![Screenshot 2022-04-14 at 21.51.18](https://global.discourse-cdn.com/julialang/original/3X/c/d/cdf8e3822e6a4041140d53c2c04808efc19dc1ea.png)

**Update**

Actually, `BLAS.set_num_threads(1)` has a similar behavior on x86 systems – it makes the code run a little bit faster and with much lesser CPU utilization. But `-t 4` has a consistent behavior there. The code runs about 3x faster than `-t 1` with the same value of BLAS threads.

On this `neoverse-n1` system running `-t 4` with `BLAS.set_num_threads(1)` is more than x10 times faster then the default `BLAS.set_num_threads(4)`.
