# Slowdown with multiple instances per node?

**URL:** <https://discourse.julialang.org/t/slowdown-with-multiple-instances-per-node/40274>\
**Category:** Julia at Scale\
**Tags:** question, parallel, mpi\
**Created:** [May 27, 2020, 4:33pm UTC](https://discourse.julialang.org/t/slowdown-with-multiple-instances-per-node/40274 "2020-05-27T16:33:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sebastian-steiner](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sebastian-steiner/32/15395_2.png) [@sebastian-steiner](https://discourse.julialang.org/u/sebastian-steiner)\
**Post date:** [May 27, 2020, 4:33pm UTC](https://discourse.julialang.org/t/slowdown-with-multiple-instances-per-node/40274/1 "2020-05-27T16:33:39Z")

</div>

I am currently comparing MPI performance between C and Julia and when I only have a single task per node, Julia’s performance is pretty much exactly the same as C’s. But as soon as I utilize all 32 cores on both sockets, I see a slowdown with a factor of ~2 compared to C. Has anybody else observed something like this?

This happens only at larger message sizes like 1024000 bytes.

P.S.: I use slurm with the parameter `--ntasks-per-node 32` to have multiple instances per node if that matters.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [May 27, 2020, 5:09pm UTC](https://discourse.julialang.org/t/slowdown-with-multiple-instances-per-node/40274/2 "2020-05-27T17:09:55Z")

</div>

Is your code using BLAS (matrix multiplication and the like)? If so, set blas threads to 1 when you are running on all cores.

---

<div class="post-metadata">

**Author:** ![sebastian-steiner](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sebastian-steiner/32/15395_2.png) [@sebastian-steiner](https://discourse.julialang.org/u/sebastian-steiner)\
**Post date:** [May 27, 2020, 5:20pm UTC](https://discourse.julialang.org/t/slowdown-with-multiple-instances-per-node/40274/3 "2020-05-27T17:20:17Z")

</div>

No, not really. I only call Allreduce in a tight loop:

```julia
send = zeros(UInt8, msize)
recv = zeros(UInt8, msize)
for i in 1:args.nrep
    # start sync
    MPI.Barrier(comm)

    times[i] = MPI.Wtime()
    MPI.Allreduce!(send, recv, msize, args.operation, comm)
    times[i] = MPI.Wtime() - times[i]
end

```

Could I still benefit from setting blas threads to 1?
