# Julia seems to be running multithreaded when I don't want it to

**URL:** <https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391>\
**Category:** Julia at Scale\
**Created:** [March 1, 2023, 2:39pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391 "2023-03-01T14:39:50Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 4:38pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/21 "2023-03-01T16:38:21Z")

</div>

You mean in the Julia script? If so, yes I did.

---

<div class="post-metadata">

**Author:** ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)\
**Post date:** [March 1, 2023, 4:50pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/22 "2023-03-01T16:50:32Z")

</div>

I took the liberty of optimizing your Julia implementation a bit. I removed all of the slicing allocations by using `@view` and used `mul!` for in-place multiplication. Neither of these changed the runtime significantly at these large problem sizes, but they’ll reduce the memory churn which is never a bad thing. More significantly, I also used `@view` to trim computations on parts of your arrays that were still all-zero for a 2x speedup.

I checked that this still gave virtually the same results, but you should check it too.

```julia
using LinearAlgebra: norm, mul!

function CGS2_alt(X::AbstractMatrix)
  Base.require_one_based_indexing(X) # ensure X has 1-based indices
  T = float(eltype(X))
  n, m = size(X)
  Q = zeros(T, n, m)
  R = zeros(T, m, m)
  r = zeros(T, m)
  w = zeros(T, n)
  R[1, 1] = norm(@view X[:, 1])
  Q[:, 1] .= @view(X[:, 1]) ./ R[1, 1]
  for i in 2:m
	iprev = 1:i-1
    w .= @view X[:, i]
	Qi = @view Q[:,iprev]
	ri = @view r[iprev]
    mul!(ri, Qi', w) # ri .= (w'Qi)'
    R[iprev, i] = ri
    mul!(w, Qi, ri, -1, 1) # w .-= Qi * ri
    mul!(ri, Qi', w) # ri .= (w'Qi)'
    mul!(w, Qi, ri, -1, 1) # w .-= Qi * ri
    R[i, i] = norm(w)
    Q[:, i] .= w ./ R[i, i]
  end
  return Q, R
end

```

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 1, 2023, 4:51pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/23 "2023-03-01T16:51:43Z")

</div>

By the way, I think the MGS is better than CGS2. Check out this code:

```julia
using LinearAlgebra

function coldot(A, j, i)
    m = size(A, 1)
    r = zero(eltype(A))
    @simd for k in 1:m
        r += A[k, i] * A[k, j]
    end
    return r; 
end

function colnorm(A, j)
    return sqrt(coldot(A, j, j)); 
end

function colsubt!(A, i, j, r)
    m = size(A, 1)
    @simd for k in 1:m
        A[k, i] -= r * A[k, j]
    end
end

function normalizecol!(A, j)
    m = size(A, 1)
    r = 1.0 / colnorm(A, j)
    @simd for k in 1:m
        A[k, j] *= r
    end
end

function mgsortho!(A)
    @show m, n = size(A)
    normalizecol!(A, 1)
    for j in 2:n 
        Base.Threads.@threads for i in j:n
            colsubt!(A, i, j-1, coldot(A, j-1, i))
        end
        normalizecol!(A, j)
    end
    return A
end

"""
Orthogonalize!(V, hv, vv, orth; verbose=false)

C. T. Kelley, 2022

Orthogonalize the Krylov vectors using your (my) choice of
methods. Anything other than classical Gram-Schmidt twice (cgs2) is
likely to become an undocumented and UNSUPPORTED option. Methods other 
than cgs2 are mostly for CI for the linear solver.

DO NOT use anything other than "cgs2" with Anderson acceleration.
"""
function Orthogonalize!(V, hv, vv, orth = "cgs2"; verbose = false)
    orthopts = ["mgs1", "mgs2", "cgs1", "cgs2"]
    orth in orthopts || error("Impossible orth spec in Orthogonalize!")
    if orth == "mgs1"
        mgs!(V, hv, vv; verbose = verbose)
    elseif orth == "mgs2"
        mgs!(V, hv, vv, "twice"; verbose = verbose)
    elseif orth == "cgs1"
        cgs!(V, hv, vv, "once"; verbose = verbose)
    else
        cgs!(V, hv, vv, "twice"; verbose = verbose)
    end
end

"""
mgs!(V, hv, vv, orth; verbose=false)
"""
function mgs!(V, hv, vv, orth = "once"; verbose = false)
    k = length(hv) - 1
    normin = norm(vv)
    #p=copy(vv)
    @views for j = 1:k
        p = vec(V[:, j])
        hv[j] = p' * vv
        vv .-= hv[j] * p
    end
    hv[k+1] = norm(vv)
    if (normin + 0.001 * hv[k+1] == normin) && (orth == "twice")
        @views for j = 1:k
            p = vec(V[:, j])
            hr = p' * vv
            hv[j] += hr
            vv .-= hr * p
        end
        hv[k+1] = norm(vv)
    end
    nv = hv[k+1]
    #
    # Watch out for happy breakdown
    #
    #if hv[k+1] != 0
    #@views vv .= vv/hv[k+1]
    (nv != 0) || (verbose && (println("breakdown in mgs1")))
    if nv != 0
        vv ./= nv
    end
end

"""
cgs!(V, hv, vv, orth="twice"; verbose=false)

Classical Gram-Schmidt.
"""
function cgs!(V, hv, vv, orth = "twice"; verbose = false)
    #
    # no BLAS. Seems faster than BLAS since 1.6 and allocates
    # far less memory.
    #
    k = length(hv)
    T = eltype(V)
    onep = T(1.0)
    zerop = T(0.0)
    @views rk = hv[1:k-1]
    pk = zeros(T, size(rk))
    qk = vv
    Qkm = V
    # Orthogonalize
    # New low allocation stuff
    mul!(rk, Qkm', qk, 1.0, 1.0)
    ### mul!(pk, Qkm', qk)
    ### rk .+= pk
    ## rk .+= Qkm' * qk
    # qk .-= Qkm * rk
    mul!(qk, Qkm, rk, -1.0, 1.0)
    if orth == "twice"
        # Orthogonalize again
        # New low allocation stuff
        mul!(pk, Qkm', qk)
        ## pk .= Qkm' * qk
        # qk .-= Qkm * pk
        mul!(qk, Qkm, pk, -1.0, 1.0)
        rk .+= pk
    end
    # Keep track of what you did.
    nqk = norm(qk)
    (nqk != 0) || (verbose && (println("breakdown in cgs")))
    (nqk > 0.0) && (qk ./= nqk)
    hv[k] = nqk
end

function qrctk!(A; orth = "cgs2")
    T = typeof(A[1, 1])
    (m, n) = size(A)
    R = zeros(T, n, n)
    @views R[1, 1] = norm(A[:, 1])
    @views A[:, 1] /= R[1, 1]
    @views for k = 2:n
        hv = vec(R[1:k, k])
        Qkm = view(A, :, 1:k-1)
        vv = vec(A[:, k])
        Orthogonalize!(Qkm, hv, vv, orth)
    end
    return (Q = A, R = R)
end

@show Threads.nthreads()

m = 200000
n = 200
A = rand(m, n)
B = deepcopy(A)

A .= B
@info "mgsortho!(A)"
@time mgsortho!(A)
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qr!(A)"
@time begin
    q = qr!(A)
    A .= Matrix(q.Q)
end
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qr!(A): broken up into two steps"
@time q = qr!(A)
@time A .= Matrix(q.Q)
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qrctk!(A; orth = \"cgs1\")"
@time qrctk!(A; orth = "cgs1")
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qrctk!(A; orth = \"cgs2\")"
@time qrctk!(A; orth = "cgs2")
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qrctk!(A; orth = \"mgs1\")"
@time qrctk!(A; orth = "mgs1")
@show norm(A'*A - LinearAlgebra.I, Inf)

A .= B
@info "qrctk!(A; orth = \"mgs2\")"
@time qrctk!(A; orth = "mgs2")
@show norm(A'*A - LinearAlgebra.I, Inf)

nothing

```

The good MGS is `mgsortho!`. `mgs!` and `cgs!` are due to C. T. Kelley [GitHub - ctkelley/SIAMFANLEquations.jl: This is a Julia package for a book project.](https://github.com/ctkelley/SIAMFANLEquations.jl/)

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 4:55pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/24 "2023-03-01T16:55:50Z")

</div>

Well, yeah, MGS is twice less the FLOP count than CGS2. I implemented MGS too, I didn’t attach it here, but if I run it on the 72-cores machines, CGS2 is actually about 4X faster! This is symptomatic of a parallel computation since the number of synchronizations is less in CGS2 than in MGS.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 1, 2023, 4:58pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/25 "2023-03-01T16:58:55Z")

</div>

Try the codes above.  
This is what I get on 63 threads:

```julia
julia> include("play.jl")
Threads.nthreads() = 63
[ Info: mgsortho!(A)
(m, n) = size(A) = (200000, 200)
  8.519595 seconds (218.53 k allocations: 13.934 MiB, 4.90% compilation time)
norm(A' * A - LinearAlgebra.I, Inf) = 3.1086244689504383e-15
[ Info: qr!(A)
  5.834709 seconds (262.59 k allocations: 624.114 MiB, 1.37% gc time, 6.05% compilation time)
norm(A' * A - LinearAlgebra.I, Inf) = 2.4424906541753444e-15
[ Info: qr!(A): broken up into two steps
  2.931786 seconds (5 allocations: 112.625 KiB)
  2.324473 seconds (9 allocations: 610.657 MiB, 6.39% gc time)
norm(A' * A - LinearAlgebra.I, Inf) = 2.4424906541753444e-15
[ Info: qrctk!(A; orth = "cgs1")
 12.793860 seconds (3.44 M allocations: 170.772 MiB, 0.08% gc time, 19.62% compilation time)
norm(A' * A - LinearAlgebra.I, Inf) = 3.720802284985433e-14
[ Info: qrctk!(A; orth = "cgs2")
 20.320473 seconds (404 allocations: 2.015 MiB)
norm(A' * A - LinearAlgebra.I, Inf) = 2.4424906541753444e-15
[ Info: qrctk!(A; orth = "mgs1")
 46.986516 seconds (40.20 k allocations: 29.656 GiB, 0.80% gc time)
norm(A' * A - LinearAlgebra.I, Inf) = 2.220446049250313e-15
[ Info: qrctk!(A; orth = "mgs2")
 46.504030 seconds (40.20 k allocations: 29.656 GiB, 0.49% gc time)
norm(A' * A - LinearAlgebra.I, Inf) = 2.220446049250313e-15
[ Info: CGS2_alt(A)
 22.561234 seconds (1.86 M allocations: 399.342 MiB, 0.17% gc time, 6.44% compilation time)
norm(A' * A - LinearAlgebra.I, Inf) = 2.4424906541753444e-15

```

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [March 1, 2023, 5:47pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/26 "2023-03-01T17:47:14Z")

</div>

You could try:  
@code\_llvm CGS2(X)

and in the output look for simd instructions. On my machine I see:

```julia
vector.body: ; preds = %vector.body, %vector.ph
; ││││ @ simdloop.jl:78 within `macro expansion`
; ││││┌ @ int.jl:87 within `+`
       %index = phi i64 [0, %vector.ph], [%index.next, %vector.body]
       %126 = getelementptr inbounds double, double* %123, i64 %index
; ││││└
; ││││ @ simdloop.jl:77 within `macro expansion` @ broadcast.jl:961
; ││││┌ @ broadcast.jl:597 within `getindex`
; │││││┌ @ broadcast.jl:642 within `_broadcast_getindex`
; ││││││┌ @ broadcast.jl:666 within `_getindex`
; │││││││┌ @ broadcast.jl:636 within `_broadcast_getindex`
; ││││││││┌ @ array.jl:924 within `getindex`
           %127 = bitcast double* %126 to <4 x double>*
           %wide.load = load <4 x double>, <4 x double>* %127, align 8
; ││││││└└└
; ││││││ @ broadcast.jl:643 within `_broadcast_getindex`
; ││││││┌ @ broadcast.jl:670 within `_broadcast_getindex_evalf`
; │││││││┌ @ float.jl:386 within `/`
          %128 = fdiv <4 x double> %wide.load, %broadcast.splat
; ││││└└└└
; ││││ @ simdloop.jl:78 within `macro expansion`
; ││││┌ @ int.jl:87 within `+`
       %129 = getelementptr inbounds double, double* %125, i64 %index

```

If you see simdloops only on the new CPU and not on the old CPU this might be an explanation for the performance difference. (SIMD: single instruction multiple data)  
It might also be that the newer CPU supports AVX512 and the older only AVX256.

SIMD commands use only one CPU core, but can still process 4 or 8 words of data in parallel.

Other possible reasons for the speed difference:

- different amount of RAM
- different cache size

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [March 1, 2023, 6:04pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/27 "2023-03-01T18:04:43Z")

</div>

Following this thread a bit, and it doesn’t seem worth testing a bunch of things until @giordano’s suggestion of using `top`, `htop`, or other resource monitor is followed. I personally use `conky` on my desktop for exactly this reason. I just start a process, look at `conky` on my desktop and immediately can tell if something is multi-threading or multi-processing. It’s a super simple test.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 1, 2023, 6:11pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/28 "2023-03-01T18:11:47Z")

</div>

Also, on Win 10/11 try ProcessExplorer.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [March 1, 2023, 6:14pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/29 "2023-03-01T18:14:34Z")

</div>

I didn’t know about ProcessExplorer. On Windows, usually I will open Task Manager, click on the “Performance” tab, right click on the CPU graph, then “Change Graph To” → “Logical Processors”.

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:17pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/30 "2023-03-01T18:17:33Z")

</div>

I cannot use conky on my machine. As I said htop is filled by processor gauges and does not display the ongoing processes. However, top does work. I ran it while running the code on the 72-cores machine. It seems as if there’s only one thread running:

 ![top](https://global.discourse-cdn.com/julialang/original/3X/9/2/927bd06217e7772b1825481ee4d07b8c5b3418e5.png)

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:18pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/31 "2023-03-01T18:18:31Z")

</div>

I’m on Linux.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 1, 2023, 6:23pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/32 "2023-03-01T18:23:42Z")

</div>

If you mean was my timing data for Linux, the answer is yes.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [March 1, 2023, 6:26pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/33 "2023-03-01T18:26:25Z")

</div>

Thanks! I’m not expert, but if its only on one thread then maybe @ufechner7 is correct that there is some SIMD magic going on instead of multi-threading?

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [March 1, 2023, 6:40pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/35 "2023-03-01T18:40:49Z")

</div>

This says the usage is at `3600%` which is like using 36 cores at 100% util, so this is definitely running with multiple threads, probably on all 36 cores of your machine (I assume you have 36 actual cores which are just hyperthreaded).

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:43pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/36 "2023-03-01T18:43:25Z")

</div>

Here is what I get with `@code_llvm CGS2(X)`:

```julia
julia> @code_llvm CGS2(X)
; @ /home/mica_20/users/nxf93431/test/julia-gram-schmidt/GramSchmidt.jl:93 within `CGS2`
define void @julia_CGS2_662([2 x {}*]* noalias nocapture sret([2 x {}*]) %0, {}* nonnull align 16 dereferenceable(40) %1) #0 {
top:
  %gcframe2097 = alloca [8 x {}*], align 16
  %gcframe2097.sub = getelementptr inbounds [8 x {}*], [8 x {}*]* %gcframe2097, i64 0, i64 0
  %2 = bitcast [8 x {}*]* %gcframe2097 to i8*
  call void @llvm.memset.p0i8.i32(i8* noundef nonnull align 16 dereferenceable(64) %2, i8 0, i32 64, i1 false)
  %3 = alloca { [1 x [1 x i64]], i64 }, align 8
  %4 = alloca [1 x [1 x i64]], align 8
  %5 = alloca [1 x [1 x i64]], align 8
  %6 = alloca [1 x [1 x i64]], align 8
  %7 = alloca [1 x [1 x i64]], align 8
  %8 = alloca { [1 x [1 x i64]], i64 }, align 8
  %9 = alloca { [1 x [1 x i64]], i64 }, align 8
  %10 = alloca [1 x [1 x i64]], align 8
  %11 = alloca [1 x [2 x i64]], align 8
  %12 = alloca [1 x [1 x i64]], align 8
  %13 = alloca { [1 x [1 x i64]], i64 }, align 8
  %14 = alloca [1 x [2 x i64]], align 8
  %15 = alloca [1 x [1 x i64]], align 8
  %16 = alloca [1 x [1 x i64]], align 8
  %17 = alloca { [1 x [1 x i64]], i64 }, align 8
  %18 = alloca [1 x [1 x i64]], align 8
  %19 = alloca [1 x [1 x i64]], align 8
  %thread_ptr = call i8* asm "movq %fs:0, $0", "=r"() #7
  %ppgcstack_i8 = getelementptr i8, i8* %thread_ptr, i64 -8
  %ppgcstack = bitcast i8* %ppgcstack_i8 to {} ****
  %pgcstack = load {} ***, {}**** %ppgcstack, align 8
; @ /home/mica_20/users/nxf93431/test/julia-gram-schmidt/GramSchmidt.jl:94 within `CGS2`
; â @ array.jl:152 within `size`
  %20 = bitcast [8 x {}*]* %gcframe2097 to i64*
   store i64 24, i64* %20, align 16
   %21 = getelementptr inbounds [8 x {}*], [8 x {}*]* %gcframe2097, i64 0, i64 1
   %22 = bitcast {} **%21 to {}***
   %23 = load {} **, {}*** %pgcstack, align 8
   store {} **%23, {}*** %22, align 8
   %24 = bitcast {} ***%pgcstack to {}***
   store {} **%gcframe2097.sub, {}*** %24, align 8
   %25 = bitcast {}* %1 to {}**
   %26 = getelementptr inbounds {}*, {}** %25, i64 3
   %27 = bitcast {}** %26 to i64*
   %28 = load i64, i64* %27, align 8
   %29 = getelementptr inbounds {}*, {}** %25, i64 4
   %30 = bitcast {}** %29 to i64*
   %31 = load i64, i64* %30, align 8
; â
; @ /home/mica_20/users/nxf93431/test/julia-gram-schmidt/GramSchmidt.jl:95 within `CGS2`
; â @ array.jl:584 within `zeros` @ array.jl:588
; ââ @ boot.jl:469 within `Array` @ boot.jl:461
    %32 = call nonnull {}* inttoptr (i64 47976924541792 to {}* ({}*, i64, i64)*)({}* inttoptr (i64 47977128374960 to {}*), i64 %28, i64 %31)
; ââ
; â @ array.jl:584 within `zeros` @ array.jl:589

```

and it keeps going, I could not copy everything… There does not seem to be a sign of simdloops.

The amounts of RAM are as follows:

On the 2-cores machine:  
15G

On the 72-cores machine:  
755G

The cache sizes are as follows:

On the 2-cores machine:  
L1: 32KB  
L2: 256KB  
L3: 35MB

On the 72-cores machine:  
L1: 32KB  
L2: 1MB  
L3: 72MB

So the RAM and cache are clearly different. If this alone impacts the Julia code so much, why doesn’t it impact the C code at all? For reminder, here is the matrix of running times:

2-cores machine:  
Julia code: 166.49 s  
C code: 89.68 s

72-cores machine  
Julia code: 35.96 s  
C code: 84.54 s

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [March 1, 2023, 6:43pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/37 "2023-03-01T18:43:51Z")

</div>

You said you have:

> [@nvenkov1](#):
>
> using LinearAlgebra: BLAS  
> BLAS.set\_num\_threads(1)

at the beginning of your script.

But are you sure you are using BLAS and not MKL?

To Check Installation:

```julia
julia> using LinearAlgebra

julia> BLAS.get_config()
LinearAlgebra.BLAS.LBTConfig
Libraries: 
└ [ILP64] libopenblas64_.0.3.13.dylib

```

```julia
julia> using MKL

julia> BLAS.get_config()
LinearAlgebra.BLAS.LBTConfig
Libraries: 
└ [ILP64] libmkl_rt.1.dylib

```

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:44pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/38 "2023-03-01T18:44:45Z")

</div>

Great! Thanks for noticing. This is the only way these results make sense. Do you know how I can prevent hyperthreading?

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [March 1, 2023, 6:46pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/39 "2023-03-01T18:46:14Z")

</div>

You should launch Julia with a single thread, which can be done by setting the `JULIA_NUM_THREADS` env variable before the launch (you can’t do this after opening), or just set it as an option when you open the REPL:

```bash
julia -t 1

```

or

```bash
julia --threads=1

```

You should also set the BLAS threads as you are already doing.  
Try the test again, and use `top` to make sure that CPU usage doesn’t go over 100%. If it does, you are probably calling something external which is using more threads.

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:46pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/40 "2023-03-01T18:46:57Z")

</div>

I did add

```julia
using LinearAlgebra: BLAS
BLAS.set_num_threads(1)

```

to my script.

I do not know whether I use OpenBLAS or MKL. I thought all Julia installations used OpenBLAS by default. How can I check this? And if I use MKL, how can I limit the number of MKL threads from Julia?

---

<div class="post-metadata">

**Author:** ![nvenkov1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nvenkov1/32/43679_2.png) [@nvenkov1](https://discourse.julialang.org/u/nvenkov1)\
**Post date:** [March 1, 2023, 6:47pm UTC](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391/41 "2023-03-01T18:47:50Z")

</div>

This was done already. Prior to running the script, I exported `JULIA_NUM_THREADS=1`.

[Previous page](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391.md?page=1)

[Next page](https://discourse.julialang.org/t/julia-seems-to-be-running-multithreaded-when-i-dont-want-it-to/95391.md?page=3)
