# Thread-safe implementation for simple counter/mean

**URL:** <https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542>\
**Category:** General Usage\
**Tags:** question, multithreading\
**Created:** [June 10, 2026, 11:57am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542 "2026-06-10T11:57:00Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 10, 2026, 11:57am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/1 "2026-06-10T11:57:00Z")

</div>

I have a callable which performs an iterative calculation, which is otherwise thread-safe. I would like to keep a statistic of the mean number of iterations that it used, and is thread-safe.

_Stylized_ code follows:

```julia
mutable struct OnlineMean
    count::Int
    mean::Float64
    OnlineMean() = new(0, 0.0)
end

function update!(om::OnlineMean, x)
    om.count += 1
    om.mean += (x - om.mean) / om.count
    nothing
end

struct MyCallable
    … # other fields
    mean_iterations::OnlineMean
end

# this function called from various tasks
function (mc::MyCallable)(z)
    y, iterations = do_the_work_and_count(mc, z) # assume this is thread-safe
    update!(mc.mean_iterations, iterations) # QUESTION: make this thread-safe
    y
end

```

Is this as simple as putting a `SpinLock` in `OnlineMean` and using it in `update!`? Or should I somehow (how?) use `Channel`s to pass in updates, or atomic operations…

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 10, 2026, 12:44pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/2 "2026-06-10T12:44:50Z")

</div>

> [@Tamas\_Papp](#):
>
> Is this as simple as putting a `SpinLock` in `OnlineMean` and using it in `update!`? Or should I somehow (how?) use `Channel`s to pass in updates, or atomic operations…

Any of the above?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [June 10, 2026, 12:55pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/3 "2026-06-10T12:55:54Z")

</div>

Spinlock is simple, and since the critical section is small, probably not too much at risk for starvation.

Channel has a lock internally, so it’s the same but with more memory overhead (and more complicated to control, since you need the one writer task that you need to maintain).

Atomics are probably optimal in terms of overhead, but they’ll be a bit tricky here since you need to update two things that can’t be written atomically at the same time due to both being 8 bytes large. Consider getting some inspiration from [Sequence counters and sequential locks — The Linux Kernel documentation](https://docs.kernel.org/locking/seqlock.html) , the technique there can probably be adapted to your usecase, depending on how frequent you read/write data.

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [June 10, 2026, 1:27pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/4 "2026-06-10T13:27:46Z")

</div>

I’d use a `ReentrantLock` here by wrapping `OnlineMean` in `Lockable`. It’s probably good enough. Spinlocks are tricky to use and I doubt they provide much benefit over a `ReentrantLock`, but they do cause issues if you change the code so it can yield while the lock is taken.

I can’t think of a way to do this atomically, unless you instead store the sum and the count instead of the count and the mean.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 11, 2026, 5:46am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/5 "2026-06-11T05:46:24Z")

</div>

> [@jakobnissen](#):
>
> unless you instead store the sum and the count instead of the count and the mean

I thought about the problem and I think I could just do that, as I limit iterations to be \<500, and practically they are around 20–100. So with an `UInt64`, I can still record 2^{64-9} \approx 10^{16} observations without overflow.

Can you please give me a hint on how to do this atomically, eg implementing `update!(::OnlineMean, ::Int)` above? I read the [Julia Atomic Manifesto](https://gist.github.com/vtjnash/11b0031f2e2a66c9c24d33e810b34ec0) but I am still not sure how to use this feature. Relying on 1.12 features is OK for me.

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [June 11, 2026, 6:22am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/6 "2026-06-11T06:22:32Z")

</div>

Below is an example of using atomics. However, updates such as `@atomic om.x += 1` is implemented with a loop in LLVM (since neither armv8 or x86 has dedicated instructions for this), so on my Mac, under contention, using atomics is much **slower** than using a lock.

Here is the atomic version:

```julia
mutable struct OnlineMean
    count::Threads.Atomic{UInt64}
    sum::Threads.Atomic{UInt64}

    OnlineMean() = new(Threads.Atomic{UInt64}(0), Threads.Atomic{UInt64}(0))
end

mean(om::OnlineMean) = om.sum[] / om.count[]

# This is thread-safe. Multiple tasks can increment om.count
# before om.sum, but that will not corrupt the data in `om`,
# as both fields will be updated correctly eventually.
function update!(om::OnlineMean, x::Integer)
    ux = UInt64(x)
    # Monotonic is enough here, since we only care about the operation
    # itself being atomic, and don't care about its ordering w.r.t other
    # operations
    Threads.atomic_add!(om.count, UInt64(1))
    Threads.atomic_add!(om.sum, ux)
    om
end

function foo()
    om = OnlineMean()

    # THis loop is thread-safe
    Threads.@threads for _ in 1:10000
        for _ in 1:1000
            obs = rand(UInt(0):UInt(100))
            update!(om, obs)
        end
    end

    # NOTE: Computing the mean IS NOT threadsafe concurrently
    # with calling `update!`, since the om.count and om.sum fields
    # may be updated at different times! However, once the loop
    # above is complete, it's in the right state
    return mean(om)
end

```

And here with a lock:

```julia
mutable struct OnlineMean2
    count::UInt64
    sum::UInt64

    OnlineMean2() = new(0, 0.0)
end

mean(om::OnlineMean2) = om.sum / om.count

function update!(om::Lockable{OnlineMean2}, x::Integer)
    ux = UInt64(x)
    @lock om begin
        om[].count += 1
        om[].sum += ux
    end
    om
end

function foo2()
    om = Lockable(OnlineMean2())

    Threads.@threads for _ in 1:10000
        for _ in 1:1000
            obs = rand(UInt(0):UInt(100))
            update!(om, obs)
        end
    end

    @lock om mean(om[])
end

```

Finally, the fastest result for this case is to have each thread update seperate counters and then merge them at the end - on my computer this is 10x faster than either. But of course, the speed there depends on how expensive instantiation of these OnlineMean objects are - in this example, they’re very cheap but that might not be the case for your real world code.

```julia-auto
function update!(om::OnlineMean2, x::Integer)
    om.count += 1
    om.sum += UInt64(x)
    om
end

function foo3()
    means = [OnlineMean2() for thread in 1:10000]
    Threads.@threads for i in eachindex(means)
        om = means[i]
        for _ in 1:1000
            obs = rand(UInt(0):UInt(100))
            update!(om, obs)
        end
    end
    om = means[1]
    for other_om in @view means[2:end]
        om.count += other_om.count
        om.sum += other_om.sum
    end
    mean(om)
end

```

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [June 11, 2026, 6:49am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/7 "2026-06-11T06:49:43Z")

</div>

A few more observations:

First, I’m really surprised the lock version is faster than the atomic version. I guess it goes to show both how good our `ReentrantLock`s are, and also how slow atomic read-modify-write is.

Also, I was wrong about codegen. On my armv8 and x86 there are indeed instructions for this.

Also, atomicrmw codegen from the `@atomic` macro was updated recently ([codegen: add a pass for late conversion of known modify ops to call atomicrmw - Pull Request #57010 - JuliaLang/julia - GitHub](https://github.com/JuliaLang/julia/pull/57010)) to be more efficient. I think this lands in 1.13. In 1.12 and before, you need to use `Threads.Atomic` instead of the more general and versatile `@atomic` macro.

Finally, and very annoyingly, codegen introspection lies about atomics ([Codegen introspection misleading for`atomic_max!` · Issue #59645 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/59645)), which makes this a little tricky to dig into.

So, my takeaway:

- `ReentrantLock` remains king, and `Lockable` is the way to use it.
- Atomics are cool but they’re tricky to get right and should only be used when benchmarking verifies it makes a difference
- Structuring your code so that it requires as little as possible thread coordination, and as much work as possible are done in a single task without inter-task communication or synchronization is best.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [June 11, 2026, 7:39am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/8 "2026-06-11T07:39:48Z")

</div>

There’s a potential numerical problem here in that the order that you add the floats together will matter. What you may want to consider is storing some of numbers into a buffer where you sort them. The main idea is that you want to add numbers of similar magnitude together first.

There’s some additional explanation here:

> <https://stackoverflow.com/questions/6699066/in-which-order-should-floats-be-added-to-get-the-most-precise-result>

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 11, 2026, 8:10am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/9 "2026-06-11T08:10:45Z")

</div>

> [@jakobnissen](#):
>
> - `ReentrantLock` remains king, and `Lockable` is the way to use it.

Thanks for all those detailed benchmarks! I still don’t understand `ReentrantLock` fuly. Suppose in our `OnlineMean2` example, Task 1 acquires a lock and starts the operation, and then Task 2 also does this before Task 1 is done. Does ReentrantLock allow this? Would this not cause data corruption? (Sorry for the silly question)

> [@jakobnissen](#):
>
> - Structuring your code so that it requires as little as possible thread coordination, and as much work as possible are done in a single task without inter-task communication or synchronization is best.

I would prefer that, and would use _task_ local values, eg from OhMyThreads.jl. But I am unsure how to collect those at the end.

---

<div class="post-metadata">

**Author:** ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)\
**Post date:** [June 11, 2026, 8:25am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/10 "2026-06-11T08:25:31Z")

</div>

Just in case you missed this: OnlineStats.jl does what you want, and seems to have at least some threaded functionality: [Big Data · OnlineStats Documentation](https://joshday.github.io/OnlineStats.jl/stable/bigdata/#Threaded-Parallelism)

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [June 11, 2026, 8:49am UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/11 "2026-06-11T08:49:41Z")

</div>

If you change the OnlineMean to

```julia-auto
mutable struct OnlineMean2
@atomic count::Int
@atomic sum::Float64
end
getMean(om::OnlineMean2) = onlineMean2.sum/onlineMean2.count

```

then you can simply use atomics, no need for locks for concurrent writers (readers may still conflict with writers, though)

The reason you get away with this is that your atomic updates to the fields commute (up to floating point rounding). Your original algorithm where you store count and mean does not have this property: Updates require count and mean to be jointly atomic, not just separately atomic.

You can also do

```julia-auto
mutable struct AtomicOnlineMean
@atomic count_mean::Tuple{Int, Float64}
end

```

and do atomic updates. This is lock-free because you’re atomically updating 16 bytes and all modern architectures can do that. Yay!

If you had to carry more values than fit into 16 bytes, then your situation starts to get more complicated.

Atomics on that use a spinlock internally, with all the issues that implies (e.g.: if your system is under load, the thread holding the spinlock can get preempted in the critical section, causing all other threads to hammer the spinlock, causing more system load, which increases the chance of lock-holders getting preempted in the critical section. I am an enemy of naive spinlocks, but don’t want to digress too much).

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 11, 2026, 1:07pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/12 "2026-06-11T13:07:36Z")

</div>

> [@mkitti](#):
>
> There’s a potential numerical problem here in that the order that you add the floats together will matter. What you may want to consider is storing some of numbers into a buffer where you sort them. The main idea is that you want to add numbers of similar magnitude together first.

Sorting them still leads to an O(n) error bound though it reduces the constant coefficient (on the bound, but not necessarily on the actual error as pointed out [by Higham](https://doi.org/10.1137/0914050)), whereas [pairwise summation](https://en.wikipedia.org/wiki/Pairwise_summation) changes the bound to O(\log n) (and is also much faster than sorting and summing, assuming you use a large base case) , and [Kahan summation](https://en.wikipedia.org/wiki/Kahan_summation_algorithm) changes it to O(1). Though in all cases you might still get a large relative error if the sum is ill-conditioned (\sum |x\_i| \gg |\sum x\_i|).

If you are storing the numbers in a buffer, you might as well call `sum`, which does pairwise summation. Or just use Kahan summation, since the extra arithmetic cost is probably dwarfed by the cost of acquiring a lock. This is assuming roundoff error is a concern at all (which it may not be… you may have plenty of precision for naive summation in your problem).

If your concern is not the _magnitude_ of the error but rather that the answer be _deterministic_ (independent of the thread scheduling), of course, then I sorting is a viable option.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 11, 2026, 1:23pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/13 "2026-06-11T13:23:52Z")

</div>

Before this topic goes off on a tangent about summing floats, I just want to note that it is all integers,

> [@Tamas\_Papp](#):
>
> a statistic of the mean number of iterations that it used

so I can keep an integer sum just as well.

---

<div class="post-metadata">

**Author:** ![adienes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adienes/32/37459_2.png) [@adienes](https://discourse.julialang.org/u/adienes)\
**Post date:** [June 11, 2026, 1:32pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/14 "2026-06-11T13:32:51Z")

</div>

fastest will probably be a local sum per thread and periodically aggregate, right?

---

<div class="post-metadata">

**Author:** ![tecosaur](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tecosaur/32/23206_2.png) [@tecosaur](https://discourse.julialang.org/u/tecosaur)\
**Post date:** [June 11, 2026, 1:39pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/15 "2026-06-11T13:39:37Z")

</div>

Having read [A Concurrency Cost Hierarchy | Performance Matters](https://travisdowns.github.io/blog/2020/07/06/concurrency-costs.html), I fully expect so.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 11, 2026, 1:49pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/16 "2026-06-11T13:49:18Z")

</div>

Possibly, I am not quite confused as the advice I have seen in the Julia community is not to think about _threads_, but _tasks_, as tasks may just be reassigned to a different thread (unless I take explicit steps against this).

I don’t know how to “periodically aggregate” in this framework. I could if I could access task-local storage for all tasks, but I don’t see an API for that.

---

<div class="post-metadata">

**Author:** ![adienes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adienes/32/37459_2.png) [@adienes](https://discourse.julialang.org/u/adienes)\
**Post date:** [June 11, 2026, 1:55pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/17 "2026-06-11T13:55:20Z")

</div>

yes, I should have said “task” not “thread.” sloppy word choice

I don’t know the details of your framework, but one random idea off the top of my head might be to do the sum under a lock, but the task holds the lock most of the time (so it doesn’t have to constantly acquire it for each summand) and then releases every `% 1000 == 0` operations to wait for aggregation

---

<div class="post-metadata">

**Author:** ![WalterMadelim](https://avatars.discourse-cdn.com/v4/letter/w/3e96dc/32.png) [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Post date:** [June 11, 2026, 2:09pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/18 "2026-06-11T14:09:22Z")

</div>

> [@Tamas\_Papp](#):
>
> `y, iterations = do_the_work_and_count(mc, z) # assume this is thread-safe`

I guess the major workload happens at this line, rather than to repeatedly update the mean value.

So if let me to implement it, I would just use a `Channel` to coordinate a vector of tasks.

As suggested in Post #7, multithreading in julia is somewhat tricky, which admits of diverse methods. It’s rare for a user to predict which one could be faster. Altogether, I think `Channel` + repeatedly tasks spawning is good. And you can monitor the progress of the update of the mean (which is exactly the spirit of online). (By comparison, the methods using Locks above probably don’t achieve this.)

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [June 11, 2026, 3:14pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/19 "2026-06-11T15:14:51Z")

</div>

When using locks (including `ReentrantLock`), if task 1 takes the lock, and task 2 attempts to take the lock before task 1 has released it, then task 2 will stall until task 1 has released the lock.

For `ReentrantLock` specifically, task 2 will enter a short loop where it will repeatedly try to take the lock (i.e. it will spin). After some number of attempts, task 2 will push itself into a queue managed by the lock, and then go to sleep. Once task 1 is done, the lock will wake up the first task waiting in the queue, if any. This hybrid approach ensures low latency if the lock is quickly released, but also avoids keeping the thread busy in a loop if the lock isn’t quickly released.

In this case, since `update!` is so fast, probably most of the time, tasks waiting for the lock will be able to acquire the lock while spinning and won’t go to sleep. That’s probably why it’s fast.

---

<div class="post-metadata">

**Author:** ![langestefan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/langestefan/32/207923_2.png) [@langestefan](https://discourse.julialang.org/u/langestefan)\
**Post date:** [June 11, 2026, 3:19pm UTC](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542/20 "2026-06-11T15:19:50Z")

</div>

> [@jakobnissen](#):
>
> - `ReentrantLock` remains king, and `Lockable` is the way to use it.

What is the benefit of using `Lockable`? I am looking into using it, but it is a julia 1.11 and up feature

[Next page](https://discourse.julialang.org/t/thread-safe-implementation-for-simple-counter-mean/137542.md?page=2)
