# Local Parallel Processing Benchmark Advice

**URL:** <https://discourse.julialang.org/t/local-parallel-processing-benchmark-advice/10519>\
**Category:** Julia at Scale\
**Created:** [April 24, 2018, 10:43pm UTC](https://discourse.julialang.org/t/local-parallel-processing-benchmark-advice/10519 "2018-04-24T22:43:17Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 24, 2018, 10:43pm UTC](https://discourse.julialang.org/t/local-parallel-processing-benchmark-advice/10519/1 "2018-04-24T22:43:17Z")

</div>

ok, I spent a few days benchmarking various versions of local multiprocessing under julia 0.6.2, with 4-core and 8-core intel processors. The results are at the rear of [http://julia.cookbook.tips/doku.php?id=parallel](http://julia.cookbook.tips/doku.php?id=parallel) . roughly speaking, Threads work wonderfully as long as the function needs very little memory. Threads can deteriorate badly when each function call needs a good chunk of memory. Threads then turn worse than single-processing, which is understandable. What is less understandable is that threads then also turn worse than pmap. I am surmising that the OS has a better scheduler for such situations than Julia-internal. So, if you need to remember one thing:

```julia
use Threads if your function and its memory easily fit into the L1 cache.

```

the next important aspect to remember is that pmap deteriorates badly (3-4 orders of magnitude) with short function calls.

```julia
Never use pmap if you have many short function calls.

```

`@parallel` is useless, but not murder—like a factor 2 rather than a factor 1,000 slower than sequential processing. Finally, `@parallel` seems good for middle-of-the-road tasks, with modest memory needs and a medium amount of overhead (functions not too many and too brief).

The number of CPU or threads choice can be tweaked, but its importance pales in comparison to the advice above. Good rules of thumb (not a perfect rule) are:

- threads: use the maximum number (lower of CPU threads and function invokcations to be run).

- workers: use the maximum number of cores plus 2 as your number of workers

/iaw

---

<div class="post-metadata">

**Author:** ![anon94023334](https://avatars.discourse-cdn.com/v4/letter/a/e274bd/32.png) [@anon94023334](https://discourse.julialang.org/u/anon94023334)\
**Post date:** [April 25, 2018, 3:21am UTC](https://discourse.julialang.org/t/local-parallel-processing-benchmark-advice/10519/2 "2018-04-25T03:21:15Z")

</div>

> [@iwelch](#):
>
> Never use pmap if you have many short function calls.

You might take a look at [Proposal: Make creation of CachingPool a default for pmap · Issue #21946 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/21946) which describes the primary issue I think you’re seeing.

---

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 26, 2018, 2:32pm UTC](https://discourse.julialang.org/t/local-parallel-processing-benchmark-advice/10519/3 "2018-04-26T14:32:14Z")

</div>

indeed. but difficult and useless to check further until this is changed. for now, it is what it is.
