# Preallocating large matrices in all cores is slower than allocating every time

**URL:** <https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261>\
**Category:** Performance\
**Tags:** distributed\
**Created:** [April 14, 2021, 9:51am UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261 "2021-04-14T09:51:10Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![kaslusimoes](https://avatars.discourse-cdn.com/v4/letter/k/b19c9b/32.png) [@kaslusimoes](https://discourse.julialang.org/u/kaslusimoes)\
**Post date:** [April 14, 2021, 9:51am UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/1 "2021-04-14T09:51:10Z")

</div>

Hi everyone,

I am trying to improve a piece of code of mine (numerical simulations for a scientific research project) to scale it up.

It essentially involves running one computationally and memory heavy function for several trials and then reducing across trials. A pseudocode example:

```julia
using Distributed
addprocs()

@everywhere hard_function(x) # allocates big matrices, does stuff with them and returns an array

results = @distributed (+) for _ = 1:K # we don't want the individual results, just their sum
    hard_function(x)
end

```

This was my initial code, which was a bit slow, so I thought the fact that I am allocating big matrices (e.g. 100 \times 10^6) for every one of the K trials would be a problem. So I decided to try something else:

```julia
@everywhere matrices = preallocate(matrices)
@everywhere hard_function!(matrices, x) # defines same function, but using preallocated matrices

results = @distributed (+) for _ = 1:K
    hard_function!(matrices, x)
end

```

My reasoning being that in this way I would allocate things only once, for every worker.

It turns out it made things slower and allocates more memory (according to `@time`) 😕

I tried looking for other sources to this discrepancy, but couldn’t find an explanation so far, so maybe thought I’d ask here and check if I’m not trying anything stupid or if my original idea was wrong from the start.

Thanks a lot!  
Lucas

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [April 14, 2021, 7:23pm UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/2 "2021-04-14T19:23:56Z")

</div>

Maybe local `matrices` are transferred to the workers and you are seeing the cost of de/serialization? Maybe you can do

```julia
@everywhere begin
    matrices = ...
    getmatrices() = matrices
end

results = @distributed (+) for _ = 1:K
    hard_function!(getmatrices(), x)
end

```

* * *

Here is a demo. Let us instrument `serialize` function first:

```julia
julia> @everywhere begin
       using Serialization: Serialization, AbstractSerializer

       struct PrintSerialize
           x
       end

       function Serialization.serialize(ser::AbstractSerializer, x::PrintSerialize)
           value = x.x
           @info "Serializing:" value
           return invoke(Serialization.serialize, Tuple{AbstractSerializer,Any}, ser, x)
       end
       end

```

Define a global and a getter function:

```julia
julia> @everywhere begin
           GLOBAL = PrintSerialize([123])
           getglobal() = GLOBAL
       end

```

If I use `getglobal` function, `GLOBAL` is not serialized

```julia
julia> @distributed (+) for x in 1:3
           @show (myid(), objectid(getglobal()))
           x
       end
      From worker 2: (myid(), objectid(getglobal())) = (2, 0x2ef8e7048e9c0500)
      From worker 2: (myid(), objectid(getglobal())) = (2, 0x2ef8e7048e9c0500)
      From worker 3: (myid(), objectid(getglobal())) = (3, 0xd5452e7af34c434c)
6

```

However, if I use `GLOBAL` variable, it is serialized

```julia
julia> @distributed (+) for x in 1:3
           @show (myid(), objectid(GLOBAL))
           x
       end
┌ Info: Serializing:
│ value =
│ 1-element Vector{Int64}:
└ 123
┌ Info: Serializing:
│ value =
│ 1-element Vector{Int64}:
└ 123
      From worker 3: (myid(), objectid(GLOBAL)) = (3, 0x2c14a9407614be5b)
      From worker 2: (myid(), objectid(GLOBAL)) = (2, 0x66a2f372e6cb38ae)
      From worker 2: (myid(), objectid(GLOBAL)) = (2, 0x66a2f372e6cb38ae)
6

```

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [April 17, 2021, 2:52am UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/3 "2021-04-17T02:52:03Z")

</div>

FYI, as of FLoops.jl v0.1.10, [the private variables created with `@init`](https://juliafolds.github.io/FLoops.jl/dev/howto/parallel/#private-variables) are not serialized when used with Distributed. So, this should just work:

```julia
@floop DistributedEx() for x in 1:K
    @init matrices = ...
    y = hard_function!(matrices, x)
    @reduce result += y
end
use(result)

```

Ref: [https://github.com/JuliaFolds/FLoops.jl/pull/80](https://github.com/JuliaFolds/FLoops.jl/pull/80)

---

<div class="post-metadata">

**Author:** ![kaslusimoes](https://avatars.discourse-cdn.com/v4/letter/k/b19c9b/32.png) [@kaslusimoes](https://discourse.julialang.org/u/kaslusimoes)\
**Post date:** [April 20, 2021, 3:13pm UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/4 "2021-04-20T15:13:31Z")

</div>

Hi @tkf,

Thank you for your suggestions. I tried both of them:

- the first one didn’t help (maybe my fault as I’m not quite sure of what is going on - I had the impression that I was already accounting for the effect of “re-sending matrices” when I preallocated them inside the worker with `@everywhere` ? what is the point of calling `getmatrices()`?)
- the second one gave me the error `LoadError: syntax: invalid syntax in "local" declaration around /home/lucas/.julia/packages/FLoops/yli1G/src/reduce.jl:615`

Sorry for the delayed reply.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [April 20, 2021, 5:24pm UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/5 "2021-04-20T17:24:44Z")

</div>

> [@kaslusimoes](#):
>
> I was already accounting for the effect of “re-sending matrices”

As I demonstrated in my first comment, it wasn’t working the way you expected. Distributed.jl tries to send the globals like `matrices` to remote workers, if the globals are mentioned in `@distributed` body. The function `getmatrices` is a way to workaround this problem; i.e., it let us not mention the globals in `@distributed`.

> [@kaslusimoes](#):
>
> LoadError: syntax: invalid syntax in “local” declaration

It’s hard to tell what went wrong without an MWE.

See also: [Please read: make it easier to help you](https://discourse.julialang.org/t/psa-make-it-easier-to-help-you/14757)

---

<div class="post-metadata">

**Author:** ![kaslusimoes](https://avatars.discourse-cdn.com/v4/letter/k/b19c9b/32.png) [@kaslusimoes](https://discourse.julialang.org/u/kaslusimoes)\
**Post date:** [April 21, 2021, 8:32am UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/6 "2021-04-21T08:32:21Z")

</div>

> [@tkf](#):
>
> It’s hard to tell what went wrong without an MWE.
> 
> See also: [PSA: make it easier to help you](https://discourse.julialang.org/t/psa-make-it-easier-to-help-you/14757)

Sorry, you’re right that I haven’t been helping myself very much…

I just referred to that error in case you knew (from experience) where it might have come from.

Regarding the MWE, my original code is long and confusing, hence why I’ve been trying to go with pseudocode, but you’re right, let me try that:

```julia
using Distributed, FLoops

addprocs(3)

@everywhere begin
    using Random
    simulate!(stats) = randn!(stats)
end 

M = 10; K = 10;
aggregated_stats = zeros(M+1, 10);

@floop DistributedEx() for i in 1:K
    @init stats = zeros(M+1, 10)

    simulate!(stats)
    @reduce(aggregated_stats += stats)
end

```

this does not give the issue I had before, but another one: `UndefVarError: #__##oninit_function#261 not defined` (please find the full stacktrace [here](https://pastebin.com/n1hEafBM))  
Even though it is a different error than before, I’m saying it because I think it might help me to understand what is going on.

Thanks for taking the time to help me.

---

<div class="post-metadata">

**Author:** ![kaslusimoes](https://avatars.discourse-cdn.com/v4/letter/k/b19c9b/32.png) [@kaslusimoes](https://discourse.julialang.org/u/kaslusimoes)\
**Post date:** [April 21, 2021, 8:35am UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/7 "2021-04-21T08:35:05Z")

</div>

> [@tkf](#):
>
> As I demonstrated in my first comment, it wasn’t working the way you expected. Distributed.jl tries to send the globals like `matrices` to remote workers, if the globals are mentioned in `@distributed` body.

Thank you, I understood that from your comment (although for some mistake of mine it didn’t translate in the same way on my code…)

My confusion comes from not understanding why that is the default behaviour. Don’t we already have `$` to signal when we want the variable to be sent over? Wouldn’t it be better to have as default just using the local variable?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [April 21, 2021, 7:28pm UTC](https://discourse.julialang.org/t/preallocating-large-matrices-in-all-cores-is-slower-than-allocating-every-time/59261/8 "2021-04-21T19:28:55Z")

</div>

> [@kaslusimoes](#):
>
> `UndefVarError: #__##oninit_function#261 not defined`

You’d need `@everywhere` to define the loop first. Also, the initialization of `aggregated_stats` is not FLoops-compatible. The easiest way is to let FLoops initialize it (`mwe`). If you want to re-use the memory for `aggregated_stats`, you can do it like `mwe2`.

```julia
using Distributed

addprocs(3)

@everywhere using FLoops

@everywhere begin
    using Random
    simulate!(stats) = randn!(stats)
end

@everywhere function mwe(; M = 10, K = 10)
    @floop DistributedEx() for i in 1:K
        @init stats = zeros(M+1, 10)

        simulate!(stats)
        @reduce(aggregated_stats += stats)
    end
    return aggregated_stats
end

@everywhere function mwe2(; M = 10, K = 10)
    @floop DistributedEx() for i in 1:K
        @init stats = zeros(M+1, 10)
        simulate!(stats)
        @reduce() do (aggregated_stats = zeros(M+1, 10); stats)
            aggregated_stats .+= stats
        end
    end
    return aggregated_stats
end

mwe()
mwe2()

```

For more information, see [Parallel loops · FLoops](https://juliafolds.github.io/FLoops.jl/dev/tutorials/parallel/) and [How to do X in parallel? · FLoops](https://juliafolds.github.io/FLoops.jl/dev/howto/parallel/)

> [@kaslusimoes](#):
>
> My confusion comes from not understanding why what that is the default behaviour. Don’t we already have `$` to signal when we want the variable to be sent over?

Yeah, I think that’s a legitimate assumption but it’s also useful to know that there are a lots of “magics” (implicit handling) in Distributed.jl when it comes to serialization. The issue we are seeing is that globals and constants referenced from closure are automatically serialized:

```julia
julia> X = 0; # no @everywhere

julia> remotecall_fetch(2) do
           @show X
       end
      From worker 2: X = 0
0

```

Magics like this are convenient when working interactively but unfortunately it makes reasoning about performance hard in cases like your code in the OP.
