# Slow parallel fractals

**URL:** <https://discourse.julialang.org/t/slow-parallel-fractals/35479>\
**Category:** Performance\
**Tags:** parallel, multithreading, fractal\
**Created:** [March 3, 2020, 7:13pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479 "2020-03-03T19:13:11Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![eaubanel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eaubanel/32/13327_2.png) [@eaubanel](https://discourse.julialang.org/u/eaubanel)\
**Post date:** [March 3, 2020, 7:13pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/1 "2020-03-03T19:13:11Z")

</div>

Trying `@threads` and `@distributed` with my favourite parallel test case. Can’t figure out why `@threads` is so slow and allocates so much memory and why `@distributed` is so slow.

```julia
function genFractal(α, n=100)
    niter = 255 # max number of iterations
    threshold = 10.0 # limit of |z|, before label as divergence
    len = 3.0 # len^2 is area of picture
    xmin, ymin = -1.5, -1.5
    ymax = ymin + len
    ax = len/n
    z::Complex = 0.0
    zinit::Complex = α > 0 ? 0.0 : 1.0+im
    count = Array{Int,2}(undef,n,n)
    @threads for j in 1:n
        cy = ymax - ax*j
        for i in 1:n
            cx = ax*i + xmin
            c = cx + im*cy
            nk = niter
            z = zinit
            for k in 1:niter
                if abs(z) < threshold
                    z = z^α + c
                else
                    nk = k-1
                    break
                end
            end
            @inbounds count[i,j] = nk
        end
    end
    return count
end

frac = genFractal(2,1000);
@time genFractal(2,1000);

```

running this without the `@threads` on my MacBook pro gives:

```julia
1.609887 seconds (6 allocations: 7.630 MiB)

```

Now with `@threads` and 2 threads:

```julia
6.242670 seconds (125.02 M allocations: 3.098 GiB, 4.51% gc time)

```

Look at the allocations and memory used! I’ve got much better results by turning that outer loop into recursive calls and using `@spawn`, getting 0.815582 seconds for 2 threads.

`@distributed` is not much better. Array `Count` is now a shared array:

```julia
count = SharedArray{Int,2}(n,n)
@sync @distributed for j in 1:n

```

Starting Julia with `-p 1` gives:

```julia
33.244954 seconds (1.26 k allocations: 67.547 KiB)

```

and starting with `-p 2` gives:

```julia
16.797049 seconds (598 allocations: 24.391 KiB)

```

So it scales, but is much slower than the sequential run.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [March 3, 2020, 7:47pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/2 "2020-03-03T19:47:05Z")

</div>

Just delete this line of code:

```julia
    z::Complex = 0.0

```

That’s forcing `z` to be boxed and shared across all threads/workers — with all the contention and overhead that it incurs.

Edit: and it’s also giving you the wrong answer because all the workers are racing to update/use the same `z`.

---

<div class="post-metadata">

**Author:** ![pixel27](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pixel27/32/8902_2.png) [@pixel27](https://discourse.julialang.org/u/pixel27)\
**Post date:** [March 3, 2020, 9:13pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/3 "2020-03-03T21:13:44Z")

</div>

Is there a way to tell when the compiler is doing this for you? I’d almost prefer the compiler to just refuse to generate the code and tell me a screwed up in my attempt to use multiple threads.

Having the compiler “do stuff” for me is nice, but trying to figure out what is slowing down the process is more troublesome to figure out since the synchronization is happening invisibly.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 3, 2020, 9:23pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/4 "2020-03-03T21:23:12Z")

</div>

When you put `z` outside of the loop like that you make the compiler think you want to keep it after the loop as well. It isn’t “wrong” and so the program should definitely not refuse to run – there are many cases where that is what you might want to do.

In reality, `z` is just a temporary variable that exists only within your loop. So, in this case, no it isn’t the right thing to do. The reason why was already mentioned.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [March 3, 2020, 9:23pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/5 "2020-03-03T21:23:50Z")

</div>

I mean, any of our performance introspection tools will flag this for you. That’s how I quickly identified it.

```julia
julia> @code_warntype genFractal(2, 1000)
...
  z::Core.Box # in big bold red
...

```

It’s nice to be able to do this when you want to — but it is easy to incur race conditions. Fortunately those race conditions are quite often performance bugs, too, making them easy to spot (as in this case).

---

<div class="post-metadata">

**Author:** ![pixel27](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pixel27/32/8902_2.png) [@pixel27](https://discourse.julialang.org/u/pixel27)\
**Post date:** [March 3, 2020, 9:27pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/6 "2020-03-03T21:27:06Z")

</div>

🙂 Okay that makes sense, thanks.

---

<div class="post-metadata">

**Author:** ![eaubanel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eaubanel/32/13327_2.png) [@eaubanel](https://discourse.julialang.org/u/eaubanel)\
**Post date:** [March 4, 2020, 12:23pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/7 "2020-03-04T12:23:06Z")

</div>

That did it, for both threads and distributed versions. That was a dumb race condition I set up 🙄, but the results hadn’t been affected (at least not the fractal, visually). Now the @threads macro with 2 threads gives:

```julia
0.818296 seconds (22 allocations: 7.631 MiB)
``

and @distribute with 2 processes gives:

```

0.815279 seconds (587 allocations: 24.156 KiB)

```julia

Thanks!
```

---

<div class="post-metadata">

**Author:** ![eaubanel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eaubanel/32/13327_2.png) [@eaubanel](https://discourse.julialang.org/u/eaubanel)\
**Post date:** [March 4, 2020, 12:25pm UTC](https://discourse.julialang.org/t/slow-parallel-fractals/35479/8 "2020-03-04T12:25:20Z")

</div>

Good to know about this. Now I’ve got to study the Performance Tips page carefully!
