# Multithreading increases memory allocations

**URL:** <https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355>\
**Category:** Performance\
**Created:** [September 28, 2023, 3:15pm UTC](https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355 "2023-09-28T15:15:11Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alberto\_Roman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alberto_roman/32/50670_2.png) [@Alberto\_Roman](https://discourse.julialang.org/u/Alberto_Roman)\
**Post date:** [September 28, 2023, 3:15pm UTC](https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355/1 "2023-09-28T15:15:11Z")

</div>

Hi there,  
I am trying to understand why the number of memory allocations increases with multi threading.  
I have the following code computing the difference of a vector (it is basically the diff function)  
‘’’  
function gr\_loop(x)  
for i=1:length(x) - 1  
x[i] = x[i + 1] - x[i];  
end  
return x  
end

function gr\_loop\_multi(x)  
Threads.@threads for i=1:length(x) - 1  
x[i] = x[i + 1] - x[i];  
end  
return x  
end

x = rand(1000000);  
using BenchmarkTools  
println(“Testing simple loop version”)  
@btime gr\_loop(x);  
println(“Testing mulithread loop version”)  
@btime gr\_loop\_multi(x);  
‘’’  
The output is  
‘’’  
Testing simple loop version  
178.019 μs (0 allocations: 0 bytes)  
Testing mulithread loop version  
52.634 μs (46 allocations: 3.97 KiB)  
‘’’  
The first question is then: why the number of allocations increases in the multithreaded version?  
However what is more puzzling is that if i increase the size of x from 1e+6 to 1e+9 the time of the multithreaded version becomes equal to the simple loop ones.  
This is the output with 1e+9 elements in x  
‘’’  
Testing simple loop version  
636.210 ms (0 allocations: 0 bytes)  
Testing mulithread loop version  
650.529 ms (50 allocations: 4.09 KiB)  
‘’’  
Which is the oppoisite of what one would expect from multithreading (ie more gain with larger vectors).  
Any idea of why this is happening? Aslo I am running the code on a macbook pro with 16 cores but using in these cases only 8.  
Thanks!

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [September 28, 2023, 3:19pm UTC](https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355/2 "2023-09-28T15:19:51Z")

</div>

Multithreading causes allocations.

```julia
julia> @btime fetch(Threads.@spawn 1 + 1)
  976.800 ns (5 allocations: 534 bytes)
2

```

---

<div class="post-metadata">

**Author:** ![hendri54](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hendri54/32/9621_2.png) [@hendri54](https://discourse.julialang.org/u/hendri54)\
**Post date:** [September 28, 2023, 5:42pm UTC](https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355/3 "2023-09-28T17:42:51Z")

</div>

> [@Alberto\_Roman](#):
>
> if i increase the size of x from 1e+6 to 1e+9 the time of the multithreaded version becomes equal to the simple loop ones.

I cannot replicate your observation (Julia 1.92 on M1 Macbook Pro; 8 threads)

```julia
julia> println("Testing simple loop version")
       @btime gr_loop($x);
       println("Testing mulithread loop version")
       @btime gr_loop_multi($x);
Testing simple loop version
  210.584 μs (0 allocations: 0 bytes)
Testing mulithread loop version
  83.875 μs (42 allocations: 4.34 KiB)

julia> x = rand(1000000 * 10);

julia> println("Testing simple loop version")
       @btime gr_loop($x);
       println("Testing mulithread loop version")
       @btime gr_loop_multi($x);
Testing simple loop version
  2.160 ms (0 allocations: 0 bytes)
Testing mulithread loop version
  903.041 μs (44 allocations: 4.41 KiB)

julia> x = rand(1_000_000 * 100);

julia> println("Testing simple loop version")
       @btime gr_loop($x);
       println("Testing mulithread loop version")
       @btime gr_loop_multi($x);
Testing simple loop version
  21.734 ms (0 allocations: 0 bytes)
Testing mulithread loop version
  9.626 ms (46 allocations: 4.47 KiB)

julia> x = rand(1_000_000 * 1000);

julia> println("Testing simple loop version")
       @btime gr_loop($x);
       println("Testing mulithread loop version")
       @btime gr_loop_multi($x);
Testing simple loop version
  216.451 ms (0 allocations: 0 bytes)
Testing mulithread loop version
  93.721 ms (49 allocations: 4.56 KiB)

```

---

<div class="post-metadata">

**Author:** ![mike.ingold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike.ingold/32/203749_2.png) [@mike.ingold](https://discourse.julialang.org/u/mike.ingold)\
**Post date:** [September 28, 2023, 5:56pm UTC](https://discourse.julialang.org/t/multithreading-increases-memory-allocations/104355/4 "2023-09-28T17:56:59Z")

</div>

> [@Alberto\_Roman](#):
>
> ```julia
> function gr_loop_multi(x)
> Threads.@threads for i=1:length(x) - 1
> x[i] = x[i + 1] - x[i];
> end
> return x
> end
> 
> ```

Side note: this code could suffer from a [race-condition](https://docs.julialang.org/en/v1/manual/multi-threading/#Data-race-freedom) that leads to possible correctness issues. You’re spawning threads to independently update the values in a given array, but execution order matters: has `x[i+1]` been updated yet? By parallelizing the loop you’re not getting any guarantees that this executes in-order or perfectly in sync.
