# Parallel launch of CUDA kernels

**URL:** <https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529>\
**Category:** GPU\
**Tags:** cuda, kernelabstractions\
**Created:** [November 12, 2024, 8:50am UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529 "2024-11-12T08:50:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![hexaeder](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexaeder/32/24403_2.png) [@hexaeder](https://discourse.julialang.org/u/hexaeder)\
**Post date:** [November 12, 2024, 8:50am UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/1 "2024-11-12T08:50:32Z")

</div>

I have a system with some kernels which are repeatatly called in a loop. My kernels operate on the same input data but in know at launchtime which kernels are safe to spawn asynchronously and which aren’t. A MWE would look something like this

```julia
using Pkg
pkg"activate --temp"
pkg"add KernelAbstractions, CUDA"

using KernelAbstractions, CUDA
CUDA.functional()

@kernel function kernel1(u, offset)
    I = @index(Global) + offset
    u[I] = u[I] + 1
end

@kernel function kernel2(u, offset)
    I = @index(Global) + offset
    u[I] = u[I] + 2
end

function loop(u)
    backend = get_backend(u)

    _kernel1 = kernel1(backend)
    _kernel1(u, 0; ndrange=1000)

    _kernel2 = kernel2(backend)
    _kernel2(u, 1000; ndrange=1000)

    KernelAbstractions.synchronize(backend)
    u
end

x = cu(ones(2000))
loop(x)

```

When I inspect this code with nsys, it looks like the kernels are launched one after another and not in parallel. Naivly I thought a kernel launch is always async until you put a `synchronize` in between. Is it possible to achieve the desired behavior?

My real world example is more complex. I want to launch some kernels in parallel, wait for all of them to finish and then launch a different set which depends on the results on the first call. It seems like CUDA Graph might do what I want, but I guess just launching two in parallel would be a required first step.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [November 12, 2024, 9:42am UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/2 "2024-11-12T09:42:33Z")

</div>

Kernel launches are asynchronous with respect to the host, but are executed in-order with respect to each other.

You can use Julia tasks to model that concurrency,

---

<div class="post-metadata">

**Author:** ![hexaeder](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexaeder/32/24403_2.png) [@hexaeder](https://discourse.julialang.org/u/hexaeder)\
**Post date:** [November 12, 2024, 2:17pm UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/3 "2024-11-12T14:17:43Z")

</div>

Wait, do you mean I can achieve on-device parallelism by launching the kernels from different Tasks on the Host? That’s my goal in the end: execute multiple kernels on the same data in parallel on the GPU…

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [November 12, 2024, 3:38pm UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/4 "2024-11-12T15:38:54Z")

</div>

> [@hexaeder](#):
>
> Wait, do you mean I can achieve on-device parallelism by launching the kernels from different Tasks on the Host?

Yeah exactly. [Tasks and threads · CUDA.jl](https://cuda.juliagpu.org/stable/usage/multitasking/)

---

<div class="post-metadata">

**Author:** ![hexaeder](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexaeder/32/24403_2.png) [@hexaeder](https://discourse.julialang.org/u/hexaeder)\
**Post date:** [November 12, 2024, 5:20pm UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/5 "2024-11-12T17:20:44Z")

</div>

Thanks for the link, this makes sense now! I guess alternatively I could create the needed number of CUDA streams manually and launching kernels on different streams from the same task/thread with possibly less overhead?

Follow-up question: Since my kernels are relatively small, I think it would be best to prepare them as a cuda graph to reduce overhead. Suppose my (limited experiments) with [CUDA.jl Graph Execution](https://cuda.juliagpu.org/stable/lib/driver/#Graph-Execution) are correct, this isn’t possible using the `@captured` macro because itseems to only capture a single stream in a linear graph A-\>B-\>C-\>… Is there an equivalent to the [c-API for building graphs manually](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#creating-a-graph-using-graph-apis)?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [November 13, 2024, 1:14pm UTC](https://discourse.julialang.org/t/parallel-launch-of-cuda-kernels/122529/6 "2024-11-13T13:14:28Z")

</div>

> [@hexaeder](#):
>
> s there an equivalent to the [c-API for building graphs manually](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#creating-a-graph-using-graph-apis)?

Yes, you should check the CUDA Driver API, all of which are available in CUDA.jl. So you can just call `CUDA.cuGraphAddKernelNode_v2`. Although it’d be nice to have higher-level abstractions, so if you have any successes here, consider creating a PR with any abstractions you create.
