# Writing to a shared buffer in multi-threaded for loop

**URL:** <https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373>\
**Category:** Performance\
**Tags:** multithreading\
**Created:** [June 1, 2026, 1:18pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373 "2026-06-01T13:18:46Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [June 1, 2026, 1:18pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/1 "2026-06-01T13:18:46Z")

</div>

I have a for loop that iterates over millions of elements, each of which are processed quickly. Elements update random parts of a shared buffer. I’m trying to parallelize this operation in a performant way. I have considered the following design approaches, but they each have problems:

1. Storing multiple copies of the shared buffer in a Channel, then reducing the buffers at the end. Unfortunately the time cost of getting and putting the buffer dominates the computation time, making this approach very slow.
2. Using a buffer per thread, and reducing at the end. Unfortunately there is [no reliable way of getting a thread index](https://juliafolds2.github.io/OhMyThreads.jl/stable/literate/tls/tls/#The-naive-(and-incorrect)-approach).
3. Creating a TaskLocalValue. Notwithstanding the fact that the buffer is large, and the number of tasks could be large too (for load-balancing), using too much memory, I don’t see a way to reduce TaskLocalValues after the for loop, since they get cleaned up.

Can anyone recommend a better solution?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 1, 2026, 1:33pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/2 "2026-06-01T13:33:14Z")

</div>

> [@user664303](#):
>
> I don’t see a way to reduce TaskLocalValues after the for loop, since they get cleaned up.

TaskLocalValues.jl are just a thin wrapper around `Base.task_local_storage`, which provides more flexibility. Using the latter, when you create the local buffer for the first time in a given task, you could store it in a thread-global array (protected by a lock) to later reduce over, i.e. something like:

```julia-auto
buf = get!(task_local_storage(), :MY_BUFFER_SYMBOL) do
    # create new buffer if it doesn't exist yet for this task
    lock(my_buffers_lock) do
        newbuf = create_new_buffer()
        push!(my_buffers, newbuf)
        newbuf
    end
end::SomeBufferType

```

Alternatively, you can replace your threaded loop with a loop over equal-sized chunks of the array, via ChunkSplitters.jl, and allocate one buffer per chunk, as described here: [No more threadid indexing? [thread-local storage] - #13 by lmiq](https://discourse.julialang.org/t/no-more-threadid-indexing-thread-local-storage/131535/13) — the tradeoff is that this pushes more load-balancing responsibility onto you.

---

<div class="post-metadata">

**Author:** ![WalterMadelim](https://avatars.discourse-cdn.com/v4/letter/w/3e96dc/32.png) [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Post date:** [June 1, 2026, 3:02pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/3 "2026-06-01T15:02:02Z")

</div>

> [@user664303](#):
>
> I have a for loop that iterates over millions of elements, each of which are processed quickly. Elements update random parts of a shared buffer

I think the description is too vague. Is it really random? Or not.

If let me have a wild guess, it sounds like constructing a LP model with the GRBaddvar api.

> **[Model Creation and Modification - Gurobi Optimizer Reference Manual](https://docs.gurobi.com/projects/optimizer/en/current/reference/c/model.html#c.GRBaddvar)**
>
> Use the functions in this section to create models, populate or modify them with variables and constraints, and finally free them again.

say, you have 1million variables, the buffer is the constraint matrix.

it probably doesn’t make sense to use multithreading in the context here.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [June 1, 2026, 3:16pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/4 "2026-06-01T15:16:07Z")

</div>

> [@WalterMadelim](#):
>
> I think the description is too vague. Is it really random? Or not.

It’s a for loop to compute the Hessian of an optimization problem. Each element is a cost function that updates various sub-blocks of the Hessian. The sub-blocks are known, so not really random, but each cost function can update multiple sub-blocks, so it isn’t possible to ensure different threads access mutually exclusive parts of the buffer - unless I substantially change the program design (which has got me thinking 😅).

> [@WalterMadelim](#):
>
> If let me have a wild guess, it sounds like constructing a LP model with the GRBaddvar api.

Interesting guess 😄

> [@WalterMadelim](#):
>
> it probably doesn’t make sense to use multithreading in the context here.

In my context, multi-threading definitely makes sense.

---

<div class="post-metadata">

**Author:** ![matthias314](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Post date:** [June 1, 2026, 9:59pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/5 "2026-06-01T21:59:58Z")

</div>

> [@stevengj](#):
>
> `task_local_storage()`

Could it be that this zero-argument method of `task_local_storage` is not documented? I only get docstrings for 1, 2 or 3 arguments.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 1, 2026, 10:44pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/6 "2026-06-01T22:44:50Z")

</div>

> [@matthias314](#):
>
> Could it be that this zero-argument method of `task_local_storage` is not documented?

> <https://github.com/JuliaLang/julia/pull/61962>
>
> I notice that \`task\_local\_storage\` is documented but its \`task\_local\_storage()\` …method, which directly returns the dictionary, is not. This is a useful method because then one can use a wider range of dictionary functions, perhaps the most useful of which is \`get!\` (to either fetch or initialize an entry).
> 
> It also seems to be worthwhile to document that this is an \`IdDict\` — and, in particular, that key lookup in the task-local dictionary is based on object identity.
> 
> (On the other hand, are there certain usages of the task-local dictionary that should be warned against? e.g. iterating over it? An alternative would be to define and export a function, e.g. \`task\_local\_storage!\`, mimics \`get!\`, another function for \`haskey\`, another for \`delete!\`, etcetera, but this seems cumbersome if many methods are desired.)

---

<div class="post-metadata">

**Author:** ![WalterMadelim](https://avatars.discourse-cdn.com/v4/letter/w/3e96dc/32.png) [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Post date:** [June 2, 2026, 12:14am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/7 "2026-06-02T00:14:03Z")

</div>

> [@user664303](#):
>
> It’s a for loop to compute the Hessian of an optimization problem.

Since Hessian is used, it’s a nonlinear optimization. Is it nonconvex?

> [@user664303](#):
>
> each cost function can update multiple sub-blocks

That doesn’t seem reasonable. Each cost function should only associate to a particular block.

Block decomposition for a large-scale optimization with separable structure is well studied. Personally I’m pretty familiar with this topic and I’ve been writing algorithms in julia with multithreading APIs. The situation you’re currently describing make me think that your method of decomposition is not very ideal.

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [June 2, 2026, 2:32am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/8 "2026-06-02T02:32:34Z")

</div>

From your problem description it seems that you could avoid contention by spawning one task to manage the buffer, and attach it to a `Channel`. Other computational tasks would then send terms and location information over the channel, preferably in groups to reduce overhead.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [June 2, 2026, 7:49am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/9 "2026-06-02T07:49:32Z")

</div>

> [@WalterMadelim](#):
>
> The situation you’re currently describing make me think that your method of decomposition is not very ideal.

Thanks for trying to help, @WalterMadelim . I’ll write to you separately about this.

I’d prefer to keep this thread limited to the problem I described, since that will be of most help to others seeking solutions to that specific problem. I’m grateful for those answers that accept the premise of the question.

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [June 2, 2026, 7:52am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/10 "2026-06-02T07:52:57Z")

</div>

> [@Ralph\_Smith](#):
>
> From your problem description it seems that you could avoid contention by spawning one task to manage the buffer, and attach it to a `Channel`. Other computational tasks would then send terms and location information over the channel, preferably in groups to reduce overhead.

I agree that’s another possible approach. However, it requires a significant change to the single threaded version, and is a very bespoke solution. I’d rather have a more generic solution that’s also simpler to implement, given the single threaded implementation.

---

<div class="post-metadata">

**Author:** ![WalterMadelim](https://avatars.discourse-cdn.com/v4/letter/w/3e96dc/32.png) [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Post date:** [June 2, 2026, 8:27am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/11 "2026-06-02T08:27:51Z")

</div>

Well, actually you can, e.g. transfer this thread to the optimization category, so others can also share opinions. (Who knows if this is an XY problem, though.) I think the real helpful mindset is to use the right design and right tools to tackle the specific problems encountered in the real world.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [June 2, 2026, 11:21am UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/12 "2026-06-02T11:21:53Z")

</div>

> [@stevengj](#):
>
> Using the latter, when you create the local buffer for the first time in a given task, you could store it in a thread-global array (protected by a lock) to later reduce over

This approach tends to sacrifice reproducibility: The order of task-local buffers stored in the global lock-protected array depends on scheduler/races, and floating point ops are not associative. That makes debugging harder (you cannot checkpoint the state after applying the updates).

> [@stevengj](#):
>
> Alternatively, you can replace your threaded loop with a loop over equal-sized chunks of the array, via [ChunkSplitters.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/ChunkSplitters), and allocate one buffer per chunk, as described here:

I would recommend this ^

---

<div class="post-metadata">

**Author:** ![langestefan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/langestefan/32/207923_2.png) [@langestefan](https://discourse.julialang.org/u/langestefan)\
**Post date:** [June 2, 2026, 2:00pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/13 "2026-06-02T14:00:08Z")

</div>

OhMyThreads.jl may be a more user-friendly way to achieve the same result. See: [Thread-Safe Storage · OhMyThreads.jl](https://juliafolds2.github.io/OhMyThreads.jl/stable/literate/tls/tls/)

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [June 2, 2026, 3:58pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/14 "2026-06-02T15:58:07Z")

</div>

> [@user664303](#):
>
> Using a buffer per thread, and reducing at the end. Unfortunately there is [no reliable way of getting a thread index](https://juliafolds2.github.io/OhMyThreads.jl/stable/literate/tls/tls/#The-naive-(and-incorrect)-approach).

It’s easy to do it manually:

```julia-auto
buffers = [makebuf() for _ in 1:nthreads()]
@sync for thrid in 1:nthreads()
    @spawn begin
        buf = buffers[thrid]
        < do stuff to buf >
    end
end
< reduce over buffers >

```

---

<div class="post-metadata">

**Author:** ![user664303](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/user664303/32/37843_2.png) [@user664303](https://discourse.julialang.org/u/user664303)\
**Post date:** [June 2, 2026, 5:09pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/15 "2026-06-02T17:09:19Z")

</div>

Am I right in thinking that either I need to have a way to get a threadsafe buffer _every iteration_, or else use an explicit chunk split across threads? The latter means no automatic load balancing, while the former means some overhead. Unless I’m mistaken, it’s impossible to avoid this trade-off.

I need to check the overhead of `get!(task_local_storage(), :MY_BUFFER_SYMBOL)` to see how much improvement this makes over `Channel`. The overhead may be sufficiently small.

---

<div class="post-metadata">

**Author:** ![sgaure](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sgaure/32/14779_2.png) [@sgaure](https://discourse.julialang.org/u/sgaure)\
**Post date:** [June 3, 2026, 4:30pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/16 "2026-06-03T16:30:28Z")

</div>

> [@user664303](#):
>
> Am I right in thinking that either I need to have a way to get a threadsafe buffer _every iteration_, or else use an explicit chunk split across threads?

One way to do it is a dynamic chunk split, i.e. take long chunks in the beginning and cutting it down. I’ve sometimes done things like this:

```julia-auto
using Base.Threads

buffers = [makebuf() for _ in 1:nthreads()]
index = Atomic{Int}(1)
@sync for thrid in 1:nthreads()
    @spawn begin
        local buf = buffers[thrid]
        local chunksize = < some starting value, like largestindex/2nthreads() or similar >
        while (firstidx = atomic_add!(index, chunksize)) <= largestindex
            for idx in firstidx:min(firstidx+chunksize-1, largestindex)
                < process index idx, store somewhere in buf >
            end
            chunksize = <compute new chunk size,e.g. as max(10, (largestindex - index[]) / 2nthreads())>
        end
    end
end
< reduce buffers >

```

In this way there’s one buffer for each thread. You process a chunk of indices at a time to avoid overhead, but the chunk size can be made smaller as you approach the last index to load balance.

---

<div class="post-metadata">

**Author:** ![droodman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/droodman/32/26652_2.png) [@droodman](https://discourse.julialang.org/u/droodman)\
**Post date:** [June 7, 2026, 3:01pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/17 "2026-06-07T15:01:19Z")

</div>

I’ve been doing it like this. Let `M` be the number of loop iterations:

```julia-auto
nt = min(M, Threads.nthreads())
batches = [floor(Int,M/nt*(t-1))+1:floor(Int,M/nt*t) for t ∈ 1:nt] # indexes to split simulations by CPU thread
...
Threads.@threads for i ∈ 1:nt
    @inbounds for j ∈ batches[i]
       ...
    end
end

```

The `j` index can then be used to read and write various arrays.

But one thing worth thinking about here is “false sharing”, when the CPU throws away data in a memory buffer because adjacent memory locations were modified. Though I understand the latest version of Julia has a feature to help with this, which I haven’t wrapped my head around yet.

If the per-`j` information that is being stored by the computation is just going to be reduced with addition, then you can increase speed by doing each batch’s reduction within the `j` loop, accumulating on the fly. Then you only have to save `nt` results, not `M` results.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 7, 2026, 5:53pm UTC](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/18 "2026-06-07T17:53:38Z")

</div>

> [@droodman](#):
>
> I’ve been doing it like this.

Seems essentially equivalent to the ChunkSplitters.jl approach linked [above](https://discourse.julialang.org/t/writing-to-a-shared-buffer-in-multi-threaded-for-loop/137373/2)? That is, you decide in advance how to divide your M tasks among the threads. Seems easier to use ChunkSplitters, though?
