# Multithreading with shared memory caches

**URL:** <https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194>\
**Category:** Performance\
**Tags:** question, performance, multithreading\
**Created:** [June 11, 2023, 5:25pm UTC](https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194 "2023-06-11T17:25:07Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![danielwe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielwe/32/35657_2.png) [@danielwe](https://discourse.julialang.org/u/danielwe)\
**Post date:** [June 11, 2023, 9:57pm UTC](https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194/6 "2023-06-11T21:57:43Z")

</div>

Thanks for this! I wasn’t aware of `Iterators.partition`.

I suppose if you want the load balancing of multiple chunks per thread without having to spawn more tasks than threads, you could combine our approaches:

```julia
chunks_per_thread = 2
chunks = Iterators.partition(eachindex(A), length(A) ÷ (Threads.nthreads() * chunks_per_thread))

ch = Channel{eltype(chunks)}(Inf)
foreach(chunk -> put!(ch, chunk), chunks)
close(ch)

tasks = map(1:Threads.nthreads()) do _
    Threads.@spawn begin
        cache = zeros(3)
        for chunk in ch
            for i in chunk
                f!(cache, A[i])
                B[i] = g(cache, A[i])
            end
        end
        return cache
    end
end

caches = fetch.(tasks)

```

For long-running compute-bound tasks, it seems plausible that the overhead from the occasional lock when taking the next chunk from the channel would be less than the overhead from tasks stepping on each other due to there being more tasks than threads. Of course, one should profile to know for sure.

---

_[View the full topic](https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194)._
