# Shared-memory parallelization with large matrix

**URL:** <https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076>\
**Category:** Performance\
**Created:** [September 23, 2019, 2:46pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076 "2019-09-23T14:46:39Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sam/32/4101_2.png) [@Sam](https://discourse.julialang.org/u/Sam)\
**Post date:** [September 23, 2019, 2:46pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/1 "2019-09-23T14:46:39Z")

</div>

Hi!

I am currently writing a program that looks something like this

```julia
Threads.@threads for i = 1:m 
    a[m] = f(args, matrix[:,:,m])
end 

```

Where `f` is a function that runs some calculations, `matrix` is a large (e.g. 3000x200x`m`) matrix, and `args` some other inputs.

The problem I have is that I don’t obtain a huge improvement in performance when running this small program with many threads (with 20 threads I only get a 2x speed-up).

I think the problem is some shared-memory issues since the threads in parallel have to read from `matrix`. Any ideas on how I can fix this problem?

My current idea is to reduce the number of allocations in `f` and use `view`’s instead of slice operations. But I am not sure if this is a good idea(?).

I also have a similar program that looks like

```julia
Threads.@threads for i = 1:m 
    a[m] = g(args)
end 

```

And for this case, I get a good improvement when using many threads (10x speed-up with 20 threads).

Thus, the problem with the first program seems to be memory issues related to reading from `matrix` in parallel.

So, I guess my questions are: How can I efficiently read parts of a matrix in parallel without introducing over-head issue? Can I declare a part of a matrix as local to one particular thread?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 23, 2019, 2:53pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/2 "2019-09-23T14:53:55Z")

</div>

did you want to ask a question about your program, possibly posting more code?

---

<div class="post-metadata">

**Author:** ![Sam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sam/32/4101_2.png) [@Sam](https://discourse.julialang.org/u/Sam)\
**Post date:** [September 23, 2019, 3:08pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/3 "2019-09-23T15:08:16Z")

</div>

Unfortunately managed to post the question before finishing writing it… But now I have edited my original question 🙂

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [September 23, 2019, 3:19pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/4 "2019-09-23T15:19:11Z")

</div>

Have you tried using `view`s instead of copying the matrix content?

---

<div class="post-metadata">

**Author:** ![dalejordan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dalejordan/32/2196_2.png) [@dalejordan](https://discourse.julialang.org/u/dalejordan)\
**Post date:** [September 23, 2019, 5:25pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/5 "2019-09-23T17:25:41Z")

</div>

Do you mean  
` a[i] = f(args, matrix[:,:,i])`  
instead of  
`a[m] = f(args, matrix[:,:,m])`  
?

---

<div class="post-metadata">

**Author:** ![Sam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sam/32/4101_2.png) [@Sam](https://discourse.julialang.org/u/Sam)\
**Post date:** [September 23, 2019, 6:28pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/6 "2019-09-23T18:28:12Z")

</div>

Yes! The loop should be

```julia
Threads.@threads for m = 1:M 
   a[m] = f(args, matrix[:,:,m])
end 

```

Thanks for finding this typo!

---

<div class="post-metadata">

**Author:** ![tkluck](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkluck/32/15769_2.png) [@tkluck](https://discourse.julialang.org/u/tkluck)\
**Post date:** [September 23, 2019, 7:48pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/7 "2019-09-23T19:48:32Z")

</div>

> [@Sam](#):
>
> The problem I have is that I don’t obtain a huge important in performance when running this small program with many threads (with 20 threads I only get a 2x speed-up).

What kind of machine are you running this on? You’ll need cpu cores to run those threads.

---

<div class="post-metadata">

**Author:** ![Sam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sam/32/4101_2.png) [@Sam](https://discourse.julialang.org/u/Sam)\
**Post date:** [September 23, 2019, 8:50pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/8 "2019-09-23T20:50:48Z")

</div>

@tkluck I am running the program on a two Intel Xeon E5-2650 v3 processors (which allow for running up to 20 tasks/threads in parallel).

---

<div class="post-metadata">

**Author:** ![cshenton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cshenton/32/9326_2.png) [@cshenton](https://discourse.julialang.org/u/cshenton)\
**Post date:** [September 24, 2019, 12:22pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/9 "2019-09-24T12:22:03Z")

</div>

Have you `export JULIA_NUM_THREADs=20` in your environment? Run `Threads.nthreads()` in your interpreter to check.

Also I would recommend setting the number of threads to the number of physical (not logical) cores for numerical workloads, otherwise you’ll have 2K threads contesting K floating point units.

In addition as another commenter mentioned, you should pass in a view, or better yet pass in the full array and an index range (since creating views allocates). For example.

```julia
using Base.Threads: @threads

function work(array, column)
    array[column] = sin.(array[column])
end

function main()
    x = rand(1000, 1000)
    
    @time @threads for i=1:1000
        work(x, i)
    end 
end

main()

```

---

<div class="post-metadata">

**Author:** ![Sam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sam/32/4101_2.png) [@Sam](https://discourse.julialang.org/u/Sam)\
**Post date:** [September 24, 2019, 6:33pm UTC](https://discourse.julialang.org/t/shared-memory-parallelization-with-large-matrix/29076/10 "2019-09-24T18:33:03Z")

</div>

@cshenton Thanks for your suggestions!

Yes, I do use `JULIA_NUM_THREADs=20` and the program is running with the correct number of threads.
