# Multi-thread speed in very large vectors

**URL:** <https://discourse.julialang.org/t/multi-thread-speed-in-very-large-vectors/125342>\
**Category:** Performance\
**Tags:** multithreading\
**Created:** [January 29, 2025, 2:01pm UTC](https://discourse.julialang.org/t/multi-thread-speed-in-very-large-vectors/125342 "2025-01-29T14:01:02Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [January 29, 2025, 3:40pm UTC](https://discourse.julialang.org/t/multi-thread-speed-in-very-large-vectors/125342/5 "2025-01-29T15:40:45Z")

</div>

Some time ago I did some experiments in this respect:

> [@Gcc vs Threads.@threads vs Threads.@spawn for large loops](https://discourse.julialang.org/t/gcc-vs-threads-threads-vs-threads-spawn-for-large-loops/34273/9):
>
> A statically scheduled loop in OpenMP splits the loop range over num\_threads and typically uses a tree to fan-out the work to the threads in parallel. At the end of the loop, the inverse of the tree is typically used for a barrier. This broadcast-barrier pair of synchronization constructs are pretty much the entire overhead for the loop and these are very well studied – each takes only hundreds to a few thousand cycles, depending on the processor and the number of threads. But what happens if i…

On a laptop quite probably can hit the memory bandwith limit. Multicore servers may have two or four or more (?) lanes to memory, depending on the task, the speedup there could be considerably larger.

---

_[View the full topic](https://discourse.julialang.org/t/multi-thread-speed-in-very-large-vectors/125342)._
