# Using Distributed: computational efficiency

**URL:** <https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570>\
**Category:** Julia at Scale\
**Created:** [June 23, 2019, 2:01pm UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570 "2019-06-23T14:01:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![FujiwaraTakumiEH](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fujiwaratakumieh/32/37975_2.png) [@FujiwaraTakumiEH](https://discourse.julialang.org/u/FujiwaraTakumiEH)\
**Post date:** [June 23, 2019, 2:01pm UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/1 "2019-06-23T14:01:32Z")

</div>

## Background:

I want to add elements to the matrix and want to use `@distributed` for acceleration,here is my code:

```julia
function fill(N)
    S = zeros(N,N)
    for i in 1:length(S)
        S[i] = i
    end  
end

```

```julia
@btime fill(1000)
  4.334 ms (2 allocations: 7.63 MiB)

```

* * *

The following is my parallel code:

```julia
using Distributed
addprocs(2)
@everywhere using SharedArrays

function fill_shared(N)
    S = SharedMatrix{Int64}(N,N)
    @sync @distributed for i in 1:length(S)
        S[i] = i
    end
end

```

```julia
@btime fill_shared(1000)
  5.458 ms (552 allocations: 23.44 KiB)

```

## Question:

Why is the code after my parallelization become longer 😂, and where is the error?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [June 24, 2019, 6:57am UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/2 "2019-06-24T06:57:30Z")

</div>

I’m not sure there’s any errors here, but simply writing 1,000 integers into a vector might not be enough to amortize the additional overhead introduced by the parallelization

---

<div class="post-metadata">

**Author:** ![alvrg](https://avatars.discourse-cdn.com/v4/letter/a/eb8c5e/32.png) [@alvrg](https://discourse.julialang.org/u/alvrg)\
**Post date:** [June 24, 2019, 8:45am UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/3 "2019-06-24T08:45:01Z")

</div>

Hi, I think there a two separate issues here. One is that the function you are benchmarking also allocates the arrays to be filled. In this case, I think that allocating a SharedArray may take a bit longer than an usual Array. The other thing, as @nilshg commented, is that there is some overhead when calling the workers, which as I read somewhere was of the order of miliseconds.

I made some benchmarks removing the array allocation to observe the overhead:

```julia
using LinearAlgebra, BenchmarkTools, Distributed, SharedArrays

function fill(S)
    for i=1:length(S)
        S[i] = i
    end  
end

function bch_fill(N)
    S = zeros(N, N)
    @btime fill($S)
    @everywhere GC.gc()
    return 
end

function dist_fill(S)
    @sync @distributed for i=1:length(S)
         S[i] = i
    end  
end

function bch_dist_fill(N)
    S = SharedArray{Int}(N, N)
    @btime dist_fill($S)
    @everywhere GC.gc()
    return 
end

# Start benchmarking.

bch_fill(1000)
bch_fill(10_000)

addprocs(2);
println("num workers: $(nworkers())")
bch_dist_fill(1000)
bch_dist_fill(10_000)

addprocs(2);
println("num workers: $(nworkers())")
bch_dist_fill(1000)
bch_dist_fill(10_000)

```

I get:

```julia
  1.504 ms (0 allocations: 0 bytes)
  154.073 ms (0 allocations: 0 bytes)

# Distributed benchmarks.
num workers: 2
  1.149 ms (303 allocations: 14.41 KiB)
  95.813 ms (311 allocations: 14.53 KiB)
num workers: 4
  1.383 ms (619 allocations: 28.63 KiB)
  71.406 ms (627 allocations: 28.75 KiB)

```

So yes, I think there is some overhead time of the order of ms if you use Distributed only for a `1000x1000` array. However, if you would want to write a larger array of size `10000x10000`, then the parallel computation would be faster.

```julia
>>> versioninfo()
Julia Version 1.1.0
Commit 80516ca202 (2019-01-21 21:24 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Xeon(R) CPU E5-2620 0 @ 2.00GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.1 (ORCJIT, sandybridge)
Environment:
  JULIA_NUM_THREADS = 12

```

---

<div class="post-metadata">

**Author:** ![FujiwaraTakumiEH](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fujiwaratakumieh/32/37975_2.png) [@FujiwaraTakumiEH](https://discourse.julialang.org/u/FujiwaraTakumiEH)\
**Post date:** [June 24, 2019, 12:42pm UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/4 "2019-06-24T12:42:38Z")

</div>

Hello, I agree with you. Here I will ask a small question, assuming that I will increase the size of the array, for example:

```julia
function fill(N)
    S = zeros(N,N)
    for i in 1:length(S)
        S[i] = i
    end  
end

@btime fill(100000)

```

It will return an error:

```julia
OutOfMemoryError()

```

Does creating an array of 100000×100000 require a lot of memory? My computer RAM is 8G,I don’t understand this very well. Is this because my code is wrong or my computer performance is not enough？😅

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [June 24, 2019, 12:46pm UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/5 "2019-06-24T12:46:10Z")

</div>

```julia
julia> sizeof(zeros(10_000, 10_000))
800000000

julia> Base.summarysize(zeros(10_000,10_000))
800000040

```

so roughly 0.8 GB for a 10,000 x 10,000 `Array{Float64, 2}` - your array is 100 times this so should be 80GB

---

<div class="post-metadata">

**Author:** ![FujiwaraTakumiEH](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fujiwaratakumieh/32/37975_2.png) [@FujiwaraTakumiEH](https://discourse.julialang.org/u/FujiwaraTakumiEH)\
**Post date:** [June 24, 2019, 12:50pm UTC](https://discourse.julialang.org/t/using-distributed-computational-efficiency/25570/6 "2019-06-24T12:50:12Z")

</div>

Thank you, I understand. 😁
