# CUDA.jl calling kernels in parallel?

**URL:** <https://discourse.julialang.org/t/cuda-jl-calling-kernels-in-parallel/133032>\
**Category:** GPU\
**Created:** [October 10, 2025, 12:28am UTC](https://discourse.julialang.org/t/cuda-jl-calling-kernels-in-parallel/133032 "2025-10-10T00:28:28Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jjgarzella](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jjgarzella/32/28410_2.png) [@jjgarzella](https://discourse.julialang.org/u/jjgarzella)\
**Post date:** [October 10, 2025, 12:28am UTC](https://discourse.julialang.org/t/cuda-jl-calling-kernels-in-parallel/133032/1 "2025-10-10T00:28:28Z")

</div>

Hi, I’m trying to use CUDA.jl and CUBLAS to implement a mod N matrix multiplication using Algorithm 1 of [this paper](https://dl.acm.org/doi/10.1145/3614408.3614411) (see the [HAL preprint](https://hal.sorbonne-universite.fr/hal-04117304v1/document)). The inner loop of the algorithm requires alternating calls to CUBLAS with a broadcast kernel call (to reduce the entries of a matrix modulo N).

When I run this on a single thread, it works and gives a nice speedup compared to existing CPU implementations.

However, if I try to run this on multiple threads, there are a bunch of lock conflicts, which seems to be coming from the fact that CUDA.jl kernels are always launched in sequence (see [here](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529), [here](https://discourse.julialang.org/t/cuda-jl-multiple-threads-to-initiate-same-cuda-algorithm/79728)).

However, as far as I can tell from [stuff written about CUDA C](https://developer.nvidia.com/blog/gpu-pro-tip-cuda-7-streams-simplify-concurrency/), CUDA streams seem to allow one to execute kernels in parallel? And the [CUDA.jl docs](https://cuda.juliagpu.org/stable/usage/multitasking/) say that each thread in a `@threads` macro is given it’s own cuda stream.

So I have two questions: first, why does CUDA.jl need to have locks for kernel launches? Second, is there a standard way to overcome that in a situation like mine?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [October 11, 2025, 5:52am UTC](https://discourse.julialang.org/t/cuda-jl-calling-kernels-in-parallel/133032/2 "2025-10-11T05:52:43Z")

</div>

> [@jjgarzella](#):
>
> [CUDA.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/CUDA) kernels are always launched in sequence

Only within a single task. Given that you’re using multiple threads, presumably using Julia tasks, each of those should have its own stream and thus allow concurrent execution.

> [@jjgarzella](#):
>
> why does [CUDA.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/CUDA) need to have locks for kernel launches

A lock is taken when looking up the function from the compilation cache, but that should be very quick, and doesn’t encompass the actual launch. So launching kernels itself doesn’t take a lock.
