# Huge CPU load when using GLM

**URL:** <https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209>\
**Category:** General Usage\
**Tags:** question, glm\
**Created:** [October 24, 2022, 7:32pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209 "2022-10-24T19:32:31Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 24, 2022, 7:32pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/1 "2022-10-24T19:32:31Z")

</div>

Hello all

I’m using GLM to calculate coefficients for logistic regression. I don’t recall this being an issue when I first developed this script, but recently, I’ve noticed that this script is generating a huge load on my server. The CPU usage goes up to 6000% in some cases as monitored through “top”. I’ve narrowed it down to one function as demonstrated in this MWE

```julia
df = DataFrame(:col1 => String[], :col2 => Bool[])

for i = 1:1000
	push!(df, [randstring(1),Bool(rand(0:1))]) 
end 
	
results = glm(@formula(col2 ~ col1), df, Binomial())

```

I presume that GLM is internally multi threading this calculation, but is there an option to limit the number of threads internally within GLM to prevent saturating the server resources?

---

<div class="post-metadata">

**Author:** ![awasserman](https://avatars.discourse-cdn.com/v4/letter/a/9de0a6/32.png) [@awasserman](https://discourse.julialang.org/u/awasserman)\
**Post date:** [October 24, 2022, 8:24pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/2 "2022-10-24T20:24:33Z")

</div>

You should be able to set the maximum number of threads by starting julia with the --threads option, for example

```nohighlight
julia --threads 4

```

or

```nohighlight
julia -t 4

```

If you had used `auto` for threads, this would use the number of local threads.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 24, 2022, 8:51pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/3 "2022-10-24T20:51:23Z")

</div>

Or set JULIA\_NUM\_THREADS environment variable in your shell. (And also JULIA\_EXCLUSIVE if you’re running on bare metal not a VM)

I have a Ryzen 5 with 6 cores and hyperthreading. I set that env var to 5 so Julia can have 5 cores and my desktop can still remain responsive

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 24, 2022, 9:00pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/4 "2022-10-24T21:00:13Z")

</div>

I already have that set as threads 1 in the session that called this and it doesn’t seem to make a difference. I still see it go up to a huge number (just reran this test and it went up to 3000%)

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 24, 2022, 9:56pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/5 "2022-10-24T21:56:47Z")

</div>

How many cores/hyperthreads on your box? Perhaps GLM is calling some BLAS routine that does threading

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 24, 2022, 10:01pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/6 "2022-10-24T22:01:27Z")

</div>

this server has 64 cores (2 sockets with 32 cores per socket) with 2 threads per core (128 total).

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 24, 2022, 10:36pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/7 "2022-10-24T22:36:27Z")

</div>

So, it doesn’t seem like there’s an actual problem? (in the sense it’s not using many more threads than cores)

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 24, 2022, 10:44pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/8 "2022-10-24T22:44:48Z")

</div>

How do you conclude that? It’s generating a huge load on the server so that’s still a problem.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 24, 2022, 10:46pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/9 "2022-10-24T22:46:27Z")

</div>

Well, it’s using ~ half the available cores… if I had 2 cores and one of them was being used by Julia, I wouldn’t be worried. If i had 10 cores and 5 of them were in use by Julia… I wouldn’t be worried…

If I have 64 cores and 32 are being used by Julia… should I be worried? I mean, maybe you want to limit it further… sure… but it’s not really a “huge” load for such a server.

In any case, if it’s BLAS that’s the issue, you can use `BLAS.set_num_threads()` to limit its thread count.

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 24, 2022, 10:49pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/10 "2022-10-24T22:49:32Z")

</div>

this is a shared resource server and this one process is taking up half the resources of the server.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 24, 2022, 10:50pm UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/11 "2022-10-24T22:50:42Z")

</div>

Well if it’s doing it for a long time… I guess that’s an issue, if it’s doing it for a second, maybe not. Try to set the number of threads for BLAS and see if that helps.

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 25, 2022, 2:08am UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/12 "2022-10-25T02:08:58Z")

</div>

> [@dlakelan](#):
>
> BLAS.set\_num\_threads()

WHere do I set this variable?

I tried this but it errors out

```julia
julia> BLAS.set_num_threads(1)
ERROR: UndefVarError: BLAS not defined

```

I tried to add the BLAS package, but there is no BLAS package to add.

```julia
(@v1.8) pkg> add BLAS
    Updating registry at `/prj/yeprd/server/julia/PKG/1.6/registries/General.toml`
ERROR: The following package names could not be resolved:
 * BLAS (not found in project, manifest or registry)

```

---

<div class="post-metadata">

**Author:** ![palday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palday/32/12640_2.png) [@palday](https://discourse.julialang.org/u/palday)\
**Post date:** [October 25, 2022, 2:26am UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/13 "2022-10-25T02:26:49Z")

</div>

BLAS is part of LinearAlgebra:

```julia
julia> using LinearAlgebra

julia> BLAS
LinearAlgebra.BLAS

julia> BLAS.get_num_threads()
4

```

GLM.jl doesn’t do any multithreading itself – any and all threading comes from the BLAS.

If you’re worried about consuming resources, you can also generate your data much more efficiently:

```julia
DataFrame(:col1 => [randstring(1) for _ in 1:1000], :col2 => rand(Bool, 1000))

```

Benchmarking shows that this is **much** faster:

```julia
julia> function f1()
       df = DataFrame(:col1 => String[], :col2 => Bool[])

       for i = 1:1000
               push!(df, [randstring(1),Bool(rand(0:1))]) 
       end 
       return df
       end
f1 (generic function with 1 method)

julia> f2() = DataFrame(:col1 => [randstring(1) for _ in 1:1000], :col2 => rand(Bool, 1000))
f2 (generic function with 1 method)

julia> @benchmark f1()
BenchmarkTools.Trial: 8983 samples with 1 evaluation.
 Range (min … max): 496.292 μs … 3.969 ms ┊ GC (min … max): 0.00% … 77.65%
 Time (median): 536.952 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 555.404 μs ± 203.561 μs ┊ GC (mean ± σ): 2.82% ± 6.39%

             ██▄ ▁▇▆▃▁                                          
  ▂▁▁▂▃▃▃▃▃▃█████▆██████▇▇▇█▇▇▆▅▄▄▃▃▃▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▄
  496 μs Histogram: frequency by time 633 μs <

 Memory estimate: 275.72 KiB, allocs estimate: 7479.

julia> @benchmark f2()
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 21.254 μs … 2.057 ms ┊ GC (min … max): 0.00% … 96.97%
 Time (median): 23.270 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 29.122 μs ± 79.844 μs ┊ GC (mean ± σ): 12.73% ± 4.60%

  ▄▇█▆▄▃▂▂▂▁▁▂▂▂ ▁▁▁▁ ▂
  ███████████████▇▇▇▆▇███████▆▇▆▇▇▇▆▆▅▅▃▅▃▃▃▅▅▃▃▁▁▃▄▁▄▅▃▄▃▄▅▅ █
  21.3 μs Histogram: log(frequency) by time 67.8 μs <

 Memory estimate: 105.62 KiB, allocs estimate: 2031.

```

---

<div class="post-metadata">

**Author:** ![jonjilla](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@jonjilla](https://discourse.julialang.org/u/jonjilla)\
**Post date:** [October 25, 2022, 2:44am UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/14 "2022-10-25T02:44:48Z")

</div>

Yes, this reduces the CPU load to 100%. Strangely, the total processing time is hardly changed. What could this be doing?

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 25, 2022, 2:52am UTC](https://discourse.julialang.org/t/huge-cpu-load-when-using-glm/89209/15 "2022-10-25T02:52:29Z")

</div>

Sometimes threading a calculation, especially a fast calculation, has so much overhead that it doesn’t help, it can even hurt.
