# Is this parallel performance correct?

**URL:** <https://discourse.julialang.org/t/is-this-parallel-performance-correct/4369>\
**Category:** General Usage\
**Tags:** performance, parallel\
**Created:** [June 20, 2017, 1:51pm UTC](https://discourse.julialang.org/t/is-this-parallel-performance-correct/4369 "2017-06-20T13:51:01Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon94023334](https://avatars.discourse-cdn.com/v4/letter/a/e274bd/32.png) [@anon94023334](https://discourse.julialang.org/u/anon94023334)\
**Post date:** [June 20, 2017, 1:51pm UTC](https://discourse.julialang.org/t/is-this-parallel-performance-correct/4369/1 "2017-06-20T13:51:02Z")

</div>

I saw some strange performance today and I’m wondering what I’m doing wrong here.

```julia
julia> using BenchmarkTools

julia> a = [1:1000000;];

julia> @btime map(log, a);
  13.449 ms (3 allocations: 7.63 MiB)

julia> addprocs(4)
4-element Array{Int64,1}:
 2
 3
 4
 5

julia> wp=CachingPool(workers())
CachingPool(Channel{Int64}(sz_max:9223372036854775807,sz_curr:4), Set([4, 2, 3, 5]), Dict{Tuple{Int64,Function},RemoteChannel}())

julia> @btime pmap(wp, log, a);
  133.079 s (119748792 allocations: 3.38 GiB)

```

I did notice that the bottleneck seemed to be the master process (it was at 90+% CPU consistently while the workers were between 30% and 40%), but I was surprised at how much slower the pmap code was. Any ideas? If I had to guess, it’s because the computation is small, and this is the result of lots of data movement between nodes, but it’d be nice to have someone confirm this.

---

<div class="post-metadata">

**Author:** ![adamslc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adamslc/32/3452_2.png) [@adamslc](https://discourse.julialang.org/u/adamslc)\
**Post date:** [June 20, 2017, 3:07pm UTC](https://discourse.julialang.org/t/is-this-parallel-performance-correct/4369/2 "2017-06-20T15:07:16Z")

</div>

You might want to use the `batch_size` keyword argument for `pmap`. It doesn’t completely remove the cost of communications, but it reduces them quite a bit. For example:

```julia
julia> using BenchmarkTools

julia> a = [1:1000000;];

julia> @btime map(log, a);
  17.034 ms (3 allocations: 7.63 MiB)

julia> addprocs(4)
4-element Array{Int64,1}:
 2
 3
 4
 5

julia> @btime pmap(log, a);
  67.329 s (93247393 allocations: 2.65 GiB)

julia> @btime pmap(log, a, batch_size=10000);
  3.189 s (7012233 allocations: 185.70 MiB)

```
