# Tutorial: Concurrency patterns for controlled parallelisms

**URL:** <https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651>\
**Category:** General Usage\
**Tags:** announcement, parallel, multithreading, tutorials, concurrency\
**Created:** [June 10, 2021, 2:18am UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651 "2021-06-10T02:18:22Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 10, 2021, 2:18am UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651/1 "2021-06-10T02:18:23Z")

</div>

Hi, I just wrote a quick tutorial on a couple of patterns using `Channel` that are handy for some parallel processing: [Concurrency patterns for controlled parallelisms](https://juliafolds.github.io/data-parallelism/tutorials/concurrency-patterns/).

> High-level data parallelism is the best starting point for writing parallel programs. However, it is sometimes required to control the parallelism in your program so that, e.g., the usages of the bounded resources like memory can be managed. This is where it is necessary to deal with [_concurrency_](https://blog.golang.org/waza-talk). Although there are a lot of concurrency primitives, `Channel` is the most versatile tool that Julia provides out-of-the-box. In this tutorial, we look at how to implement simple and useful patterns based on `Channel`. Some of these patterns are known as task-parallel [_algorithmic skeletons_](https://en.wikipedia.org/wiki/Algorithmic_skeleton) (or [_parallel skeletons_](https://link.springer.com/referenceworkentry/10.1007%2F978-0-387-09766-4_24)).
> 
> 1. Worker pool
> 1. Re-distribution hacks
> 
> 2. Task farm
> 3. Pipeline
> 4. Promise (request-response)
> 1. Improving the API
> 2. Wrapping single-thread API

I wrote this mainly because I couldn’t find a place this kind of usage is summarized (other than JuliaActors and Nathan Daly’s CspExamples.jl [CSP examples in Julia?](https://discourse.julialang.org/t/csp-examples-in-julia/20011) ; so let me know if you know other related resources). I think these are very common patterns used in Go community (minus `select`) and I learned these patterns by looking into their discussions.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [June 14, 2021, 7:03pm UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651/2 "2021-06-14T19:03:10Z")

</div>

Thank you for your tutorial, it is really great.

But, it is somewhat difficult to read, I’ll try to describe my issues, maybe it will help to improve it.

1. in the first example, an argument of the `allocation` function is `body`. This is very confusing, because it is a function, which is usually denoted with verb. Maybe it’s better to use word like `process` or something similar?

```julia
    function allocate(process)
        open("/dev/urandom") do file
            buffer = Vector{UInt8}(undef, buffer_length)
            process((file, buffer))
        end
    end

```

1. To be honest, worker pool construction which consists of three interwined closures is very hard to understand. It took me some time to understand, that first example can be rewritten as (I hope I am not mistaken and this is correct representation of the original idea)

```julia
let buffer_length = 2^10
    ntasks = Threads.nthreads()
    @sync for _ in 1:ntasks
        @spawn begin
            open("/dev/urandom") do file
                buffer = Vector{UInt8}(undef, buffer_length)
                for input in works
                    read!(file, buffer)
                    input[] = sum(buffer; init = 0.0)
                end
            end
        end
    end

    sum(results) / (length(results) * buffer_length)
end

```

Maybe it makes sense to add something like `Explanation` subsection, where worker pool pattern can be shown as in the snippet above and than step by step it can be transformed to the three functions version.

1. Am I correct, that if you do not need to allocate any resources, you can just use `identity` instead of `allocate`, i.e.

```julia
    workerpool(identity, works) do ref
        ref[] = sum(rand(100))
    end

```

(as I have said, it’s not easy to correctly trace all variables and functions).

1. I’ve tried to trace threads usage with the following modification

```julia
 function workerpool(work!, allocate, request; ntasks = Threads.nthreads())
    @sync for _ in 1:ntasks
        @spawn allocate() do resource
            cnt = 0
            for input in request
                cnt += 1
                work!(input, resource)
            end
            @info Threads.threadid() cnt
        end
    end
end

```

and result was

```julia
┌ Info: 4
└ cnt = 1
┌ Info: 3
└ cnt = 1
┌ Info: 2
└ cnt = 1
┌ Info: 1
└ cnt = 29

```

so it looks like all work was done mostly by one thread. Is it a problem of a scheduler or in this particular example execution time was so small, that Julia just didn’t have time to parallelize task? Or it is the problem discussed in the next section “Re-distribution hacks”?

1. In Re-distribution hacks it is written “use of `@async` impedes migration of the tasks across OS threads (which is not implemented as of Julia 1.6 but is likely to be implemented in the future Julia versions)”. What does that mean? That using `@async` is not recommended because it schedule everything on one thread? But this behaviour may change in the future and this is what is said in brackets?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 15, 2021, 9:19pm UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651/3 "2021-06-15T21:19:14Z")

</div>

Thanks for the feedbacks!

> [@Skoffer](#):
>
> in the first example, an argument of the `allocation` function is `body` . This is very confusing, because it is a function, which is usually denoted with verb.

Yes, `process` seems to be a good name here.

(Though I’m not entirely sure if this principle always applies to argument names for callbacks; e.g., is `f` in `open(f, name)` a verb?)

> [@Skoffer](#):
>
> worker pool construction which consists of three interwined closures

Ah, yes, thanks for pointing it out. I know that it’s always to better to start with concrete examples (especially in tutorial) but generalization is always attractive… I agree starting with a concrete code like you posted would be a good idea.

> [@Skoffer](#):
>
> ```julia
> workerpool(identity, works) do ref
> ref[] = sum(rand(100))
> end
> 
> ```

You’d need

```julia
workerpool(f -> f(nothing), works) do ref, _
    ref[] = sum(rand(100))
end

```

This probably should exist as an arity overload `workerpool(work!, request)`.

> [@Skoffer](#):
>
> Is it a problem of a scheduler or in this particular example execution time was so small, that Julia just didn’t have time to parallelize task?

It may be the case, but I’d say the scheduler is doing the right thing here because we are mostly(?) only doing I/O and so it’s not very meaningful to use multiple CPUs. Maybe we need more CPU-intensive works for demonstrating parallelization.

> [@Skoffer](#):
>
> That using `@async` is not recommended because it schedule everything on one thread? But this behaviour may change in the future and this is what is said in brackets?

Using `@async` is not recommended because Julia does not move the tasks across thread. If you start them from different OS threads, multiple `@async` can use multiple OS threads. The problem is that `@async` tasks will “stick” to one OS thread (Note: this is an implementation detail AFAICT but my hunch is that this behavior will not be changed at least in 1.x time-frame).

As of Julia 1.6, `@spawn` tasks also is not migrated anyway and so mentioning this is kind of unnecessary. But if you play with Julia 1.8-DEV, it now does migrate `@spawn` tasks.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [June 20, 2021, 11:56am UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651/4 "2021-06-20T11:56:40Z")

</div>

![image](https://global.discourse-cdn.com/julialang/original/3X/0/b/0bb7161450bde6329cbaba5a4b1d62bcb63ccda6.png)

this looks like `x7` is piped into the worker `f` and result `f(x4)` comes out at the other end.

I’m guessing this is meant to say `x7` is waiting to be processed next while `f(x4)` is on its way out?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 20, 2021, 8:23pm UTC](https://discourse.julialang.org/t/tutorial-concurrency-patterns-for-controlled-parallelisms/62651/5 "2021-06-20T20:23:52Z")

</div>

Yes, the latter is what I meant. Maybe it’s clearer if I don’t put `x`s before `f`s.
