# Parallel assembly of a finite element sparse matrix

**URL:** <https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947>\
**Category:** General Usage\
**Tags:** parallel\
**Created:** [March 12, 2023, 2:11am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947 "2023-03-12T02:11:37Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 12, 2023, 2:11am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/1 "2023-03-12T02:11:37Z")

</div>

There was a thread recently ([Behavior of threads](https://discourse.julialang.org/t/behavior-of-threads/95769)) concerning a problem of interest to me: how to teach a serial finite element code to perform the assembly of a global sparse matrix in parallel (on shared memory computers).

There were several good ideas (many thanks!), which in the end served as a basis of the implementation demonstrated here on the [problem of heat conduction](https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl/blob/4575b4fd6764af1c17a8ebd4fa7700cfa4237e4e/examples/steady_state/3-d/Poisson_examples.jl#L342).

Some salient observations:

1. The serial code required no modifications.
2. The implementation is based on tasks. There was no need to know which task was running the code.
3. A handful of serial pieces of computation remain, which of course limits the available parallel speed up. In particular, the construction of the sparse matrix from a triplet of vectors does not run in parallel.
4. There is also a puzzle there: Sometimes a few of the tasks take quite a bit longer than the others. For instance, running with sixteen tasks, 5 tasks take ~11 seconds each, the other 11 tasks take over 16 seconds each. This on a machine with 64 cores and plenty of memory. And, of course, perfectly balanced load.

Any ideas on how to address the serial parts of the algorithm would be appreciated, and collaboration to implement the most promising leads would be welcome.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [March 12, 2023, 6:35am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/2 "2023-03-12T06:35:15Z")

</div>

One general thing: Did you pin your Julia threads (i.e. use something like `JULIA_EXCLUSIVE=1` or ThreadPinning.jl)? If not, does it make a difference?

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [March 12, 2023, 6:37am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/3 "2023-03-12T06:37:09Z")

</div>

> [@PetrKryslUCSD](#):
>
> This on a machine with 64 cores and plenty of memory.

What CPU(s) is it exactly?

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [March 12, 2023, 6:39am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/4 "2023-03-12T06:39:46Z")

</div>

Btw, a more minimal example would be helpful. A number of the “range splitting” functions are not defined in the linked file, afaics.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 12, 2023, 6:44pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/5 "2023-03-12T18:44:54Z")

</div>

The complete example could be run as

```julia
if !isdir("FinEtoolsHeatDiff.jl")
    run(`git clone https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl.git`)
end
cd("FinEtoolsHeatDiff.jl/examples")
println("Current folder: $(pwd())")
Pkg.activate(".")
Pkg.instantiate()
@show Threads.nthreads()
include(joinpath(pwd(), "steady_state/3-d/Poisson_examples.jl"))
using .Main.Poisson_examples; Main.Poisson_examples.Poisson_FE_H20_parass_example()
using .Main.Poisson_examples; Main.Poisson_examples.Poisson_FE_H20_parass_example()

```

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 12, 2023, 6:51pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/6 "2023-03-12T18:51:24Z")

</div>

No, I did not.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 12, 2023, 6:54pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/7 "2023-03-12T18:54:26Z")

</div>

There must be something wrong with that particular hardware. It was

“4 socket(s), 64 cores, AMD Opteron™ Processor 6380”,  
“RAM = 252 GB, Cache: L3 = 6 MB L2 = 2048kiB L1 = 16 KB”

The computation scaled much better on this machine:

2x Intel Xeon E5 2670  
Cores 8  
Code Name Sandy Bridge-EP/EX  
Package Socket 2011 LGA  
Technology 32nm  
Specification Intel Xeon CPU E5-2670 0 @ 2.60GHz  
L1 Data Cache Size 8 x 32 KBytes  
L1 Instructions Cache Size 8 x 32 KBytes  
L2 Unified Cache Size 8 x 256 KBytes  
L3 Unified Cache Size 20480 KBytes  
255 GB DDR3

All tasks were running at the same wall-clock time for the second machine (which ran WSL2, btw). The part that ran in the tasks took 52.2 sec on 1 task, 26.4 on 2 tasks, 13.2 on 4 tasks, and 7.8 sec on 8 tasks.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 13, 2023, 8:24pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/8 "2023-03-13T20:24:39Z")

</div>

Alas, I spoke too soon. The other machine display similar random delays of individual tasks. Here is an example. The computation was done twice with 4 tasks. First time, all tasks ran in the same time; in the second case one task (# 3) was significantly delayed.

```julia
Threads.nthreads() = 4
[ Info: All examples may be executed with
using .Main.Poisson_examples; Main.Poisson_examples.allrun()

First run:
Mesh generation
1000000 elements
Searching nodes for BC
Number of free degrees of freedom: 3910599
Conductivity
LinearAlgebra.BLAS.get_num_threads() = 1
[ Info: All done serial 73.12100005149841
[ Info: 1: Before conductivity 2.9690001010894775
[ Info: 3: Before conductivity 2.9690001010894775
[ Info: 4: Before conductivity 2.9850001335144043
[ Info: 2: Before conductivity 3.0
[ Info: 2: After conductivity 15.765000104904175
[ Info: 1: After conductivity 15.781000137329102
[ Info: 3: After conductivity 15.812000036239624
[ Info: 4: After conductivity 15.858999967575073
[ Info: After sync 15.858999967575073
[ Info: All done 38.342000007629395

Second run:
Mesh generation
1000000 elements
Searching nodes for BC
Number of free degrees of freedom: 3910599
Conductivity
LinearAlgebra.BLAS.get_num_threads() = 1
[ Info: All done serial 73.49599981307983
[ Info: 2: Before conductivity 2.8430001735687256
[ Info: 1: Before conductivity 2.8430001735687256
[ Info: 4: Before conductivity 2.8430001735687256
[ Info: 3: Before conductivity 2.8430001735687256
[ Info: 4: After conductivity 15.28000020980835
[ Info: 2: After conductivity 15.343000173568726
[ Info: 1: After conductivity 15.343000173568726
[ Info: 3: After conductivity 28.076000213623047
[ Info: After sync 28.076000213623047
[ Info: All done 51.075000047683716

```

What is happening? Perhaps the GC kicks in in there for some reason? Or is the scheduler not working as well as it should?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 1:04am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/9 "2023-03-14T01:04:40Z")

</div>

It looks like the GC is not at fault here: I turned it off before the loop over the tasks  
(and back on after the loop): the hanging task is still there, showing up randomly.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [March 14, 2023, 5:59am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/10 "2023-03-14T05:59:52Z")

</div>

Again, what happens if you do? Are the timings more stable?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 2:52pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/11 "2023-03-14T14:52:08Z")

</div>

> [@carstenbauer](#):
>
> JULIA\_EXCLUSIVE=1

Sorry, it wasn’t clear that that was a request to try it. Will do.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 3:28pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/12 "2023-03-14T15:28:53Z")

</div>

Worse. Now I haven’t seen equal times for the tasks even once: There are always two groups of tasks, the slow ones and the fast ones.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [March 14, 2023, 3:58pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/13 "2023-03-14T15:58:49Z")

</div>

Just as a reference, [Threaded Assembly · Ferrite.jl](https://ferrite-fem.github.io/Ferrite.jl/stable/examples/threaded_assembly/) has an example of threaded assembly using Ferrite.jl. It uses a coloring algorithm (where no elements of the same color share dofs) so that it can run the assembly in parallel as well.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 5:15pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/14 "2023-03-14T17:15:05Z")

</div>

Thank you, Kristoffer. I am aware.

We approached the problem from different viewpoints. In your case the serial part is front-loaded (the pattern creation), and in my case the computation is partitioned, with no data races, but the triplet then needs to be converted into a CSC matrix (serial part).

I noticed that you get a similar sort of parallel scaling that I do. Which is good. But I wonder how you measured it, and whether you can also see the irregularities in the thread consumption of wall-clock time?

---

<div class="post-metadata">

**Author:** ![koehlerson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/koehlerson/32/13108_2.png) [@koehlerson](https://discourse.julialang.org/u/koehlerson)\
**Post date:** [March 14, 2023, 5:33pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/15 "2023-03-14T17:33:55Z")

</div>

In this issue are a lot of timings for threaded assembly with Ferrite.jl: [Threaded Assembly Performance Degradation · Issue #526 · Ferrite-FEM/Ferrite.jl · GitHub](https://github.com/Ferrite-FEM/Ferrite.jl/issues/526) Take a look especially here: [Threaded Assembly Performance Degradation · Issue #526 · Ferrite-FEM/Ferrite.jl · GitHub](https://github.com/Ferrite-FEM/Ferrite.jl/issues/526#issuecomment-1433603603)

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 6:40pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/16 "2023-03-14T18:40:01Z")

</div>

Below is the information on how to run the example. I would be very much interested in measurements performed on another machine(s).

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 14, 2023, 9:51pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/17 "2023-03-14T21:51:12Z")

</div>

Alas, LinuxPerf.jl appears to be broken. Edit: Looks like it is not available for WSL2.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 16, 2023, 2:39am UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/18 "2023-03-16T02:39:21Z")

</div>

Interesting! I implemented an alternative way based on threads. This works quite well, at least the irregularities (slow tasks mixed in with fast tasks) are gone: all threads run in the same time.

Here is the parallel speed up for the thread implementation:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/8/7/874578795554bfd5d03b35f597145f331f7e2bef.png)  
Note well that this is only the computation of the conductivity matrix and its assembly into the COO format.  
Because the conversion into the CSC format is not parallelized, the overall speed up is quite a bit worse.

Machine used in the above graph:

2x Intel Xeon E5 2670  
Cores 8  
Code Name Sandy Bridge-EP/EX  
Package Socket 2011 LGA  
Technology 32nm  
Specification Intel Xeon CPU E5-2670 0 @ 2.60GHz  
L1 Data Cache Size 8 x 32 KBytes  
L1 Instructions Cache Size 8 x 32 KBytes  
L2 Unified Cache Size 8 x 256 KBytes  
L3 Unified Cache Size 20480 KBytes  
255 GB DDR3

This machine was running WSL2 under Windows 10.

Linear heat conduction problem. 343000 serendipity quadratic elements with 3x3x3 Gauss quadrature.

Open question: what is wrong with the task-based implementation? Why are some tasks much slower than others in the same batch?

References:  
The task loop: [https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl/blob/d041cd06035547e7bdb1422a94daf006594f1393/examples/steady\_state/3-d/Poisson\_examples.jl#L336](https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl/blob/d041cd06035547e7bdb1422a94daf006594f1393/examples/steady_state/3-d/Poisson_examples.jl#L336)  
The thread loop: [https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl/blob/d041cd06035547e7bdb1422a94daf006594f1393/examples/steady\_state/3-d/Poisson\_examples.jl#L479](https://github.com/PetrKryslUCSD/FinEtoolsHeatDiff.jl/blob/d041cd06035547e7bdb1422a94daf006594f1393/examples/steady_state/3-d/Poisson_examples.jl#L479)

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [March 16, 2023, 4:15pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/19 "2023-03-16T16:15:58Z")

</div>

Further data for the same problem. This machine:  
4 socket(s), 64 cores, AMD Opteron™ Processor 6380,  
RAM = 252 GB, Cache: L3 = 6 MB L2 = 2048kiB L1 = 16 KB

 ![image](https://global.discourse-cdn.com/julialang/original/3X/4/f/4f6336ebf13cbc923da231b7c6d5111b673af3a6.png)

Again, the thread-based loop.

Edit: Extended to more threads.

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/a/ea0b9d82714841f10ea0a4694542a55a79f7db6b.png)

---

<div class="post-metadata">

**Author:** ![purplishrock](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/purplishrock/32/13451_2.png) [@purplishrock](https://discourse.julialang.org/u/purplishrock)\
**Post date:** [March 16, 2023, 7:06pm UTC](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947/20 "2023-03-16T19:06:58Z")

</div>

Am i reading the results correctly , i.e. the Xeon is giving worse multi threaded performance ?

[Next page](https://discourse.julialang.org/t/parallel-assembly-of-a-finite-element-sparse-matrix/95947.md?page=2)
