# Clarifying expected behavior of dynamic CUDA kernels

**URL:** <https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689>\
**Category:** GPU\
**Tags:** question, parallel, cuda, dynamic-parallelism\
**Created:** [January 12, 2025, 4:27am UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689 "2025-01-12T04:27:46Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![NonDairyNeutrino](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nondairyneutrino/32/221496_2.png) [@NonDairyNeutrino](https://discourse.julialang.org/u/NonDairyNeutrino)\
**Post date:** [January 12, 2025, 4:27am UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689/1 "2025-01-12T04:27:46Z")

</div>

# Dynamic Parallelism

I’ve just learned about [_Dynamic Parallelism_](https://cuda.juliagpu.org/stable/development/kernel/#Dynamic-parallelism) that is

> …useful for recursive algorithms, or for algorithms that otherwise need to dynamically spawn new work.

So consider the following example:

```julia-auto
using CUDA

threadCount = 10

function innerKernel(outerThread)
    innerThread = threadIdx().x
    @cuprintln("outer: $outerThread, inner: $innerThread")
    return
end

function outerKernelDynamic()
    outerThread = threadIdx().x
    @cuda dynamic=true innerKernel(outerThread)
    return
end

@cuda threads=threadCount outerKernelDynamic()

```

which produces

> outer: 1, inner: 1  
> outer: 2, inner: 1  
> outer: 3, inner: 1  
> outer: 4, inner: 1  
> outer: 5, inner: 1  
> outer: 6, inner: 1  
> outer: 7, inner: 1  
> outer: 8, inner: 1  
> outer: 9, inner: 1  
> outer: 10, inner: 1

# Confusion

But I’m confused what the dynamic call is actually doing. It appears to be invoking all inner calls on thread 1, but my expectation is that each inner call would be executed on different threads e.g.

> outer: 2, inner: 9  
> outer: 4, inner: 7  
> …

The example in the documentation doesn’t depend on which thread it’s running so is rather unhelpful in this case.

# Full Code

```julia-auto
using CUDA

threadCount = 10

function innerKernel(outerThread)
    innerThread = threadIdx().x
    @cuprintln("outer: $outerThread, inner: $innerThread")
    return
end

function outerKernelStatic()
    outerThread = threadIdx().x
    innerKernel(outerThread)
    return
end

function outerKernelDynamic()
    outerThread = threadIdx().x
    @cuda dynamic=true innerKernel(outerThread)
    return
end

println("Executing static kernel")
CUDA.@sync @cuda threads=threadCount outerKernelStatic()

println("Executing dynamic kernel")
CUDA.@sync @cuda threads=threadCount outerKernelDynamic()

```

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 12, 2025, 7:26am UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689/3 "2025-01-12T07:26:52Z")

</div>

> [@NonDairyNeutrino](#):
>
> It appears to be invoking all inner calls on thread 1, but my expectation is that each inner call would be executed on different threads e.g.

That’s the wrong understanding. Every dynamic kernel launch is just that, another kernel launch, so the inner kernel starts with a fresh grid where threads are numbered from 1 again.

---

<div class="post-metadata">

**Author:** ![NonDairyNeutrino](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nondairyneutrino/32/221496_2.png) [@NonDairyNeutrino](https://discourse.julialang.org/u/NonDairyNeutrino)\
**Post date:** [January 12, 2025, 5:55pm UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689/4 "2025-01-12T17:55:11Z")

</div>

I guess that…makes perfect sense. But then if the `outerKernelDynamic` is changed to launch on multiple threads as usual as

```julia
function outerKernelDynamic()
    outerThread = threadIdx().x
    @cuda dynamic=true threads=threadCount innerKernel(outerThread)
    return
end

```

there’s the big wall of an error

> ERROR: LoadError: GPUCompiler.InvalidIRError(GPUCompiler.CompilerJob{GPUCompiler.PTXCompilerTarget, CUDA.CUDACompilerParams}(MethodInstance for outerKernelDynamic(), GPUCompiler.CompilerConfig{GPUCompiler.PTXCompilerTarget, CUDA.CUDACompilerParams}(GPUCompiler.PTXCompilerTarget(v"8.6.0", v"7.8.0", true, nothing, nothing, nothing, nothing, false, nothing, nothing), CUDA.CUDACompilerParams(v"8.6.0", v"8.5.0"), true, nothing, :specfunc, false, 2), 0x0000000000006897), Tuple{String, Vector{Base.StackTraces.StackFrame}, Any}[(“call to an unknown function”, [macro expansion at execution.jl:96, outerKernelDynamic at dynamic\_kernel\_test.jl:19], “jl\_f\_tuple”), (“call to an unknown function”, [NamedTuple at boot.jl:727, macro expansion at execution.jl:96, outerKernelDynamic at dynamic\_kernel\_test.jl:19], “jl\_f\_apply\_type”), (“call to an unknown function”, [NamedTuple at boot.jl:727, macro expansion at execution.jl:96, outerKernelDynamic at dynamic\_kernel\_test.jl:19], “ijl\_new\_structv”), (“dynamic function invocation”, [macro expansion at execution.jl:96, outerKernelDynamic at dynamic\_kernel\_test.jl:19], Core.kwcall)])

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 12, 2025, 6:47pm UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689/5 "2025-01-12T18:47:42Z")

</div>

> [@NonDairyNeutrino](#):
>
> `threads=threadCount`

`threadCount` is not defined in your kernel, as should be shown by the error (`Reason: unsupported use of an undefined name (use of 'threadCount')`). You could either forward that as an arg, or look up the block size.

---

<div class="post-metadata">

**Author:** ![NonDairyNeutrino](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nondairyneutrino/32/221496_2.png) [@NonDairyNeutrino](https://discourse.julialang.org/u/NonDairyNeutrino)\
**Post date:** [January 12, 2025, 8:33pm UTC](https://discourse.julialang.org/t/clarifying-expected-behavior-of-dynamic-cuda-kernels/124689/6 "2025-01-12T20:33:18Z")

</div>

On further thought, this makes perfect sense (I think?). Basically, it’s because there is no variable `threadCount` on the device as it’s not automatically inferred or copied from the host process/shared memory to the device process/shared memory?

I find it interesting that while there’s the distinction between host and device _arrays_ with `Array` being called from and stored on the host, and `CuArray` that’s called from the host and stored on the device, and `CuDeviceArray` which is called from and stored on the device (as far as I understand), there’s seemingly no distinction between host and device _scalars_ as, say. `Float32` that’s called from and stored on the host, and likewise `CuFloat32` and `CuDevice32` to communicate that these are stored and otherwise accessible to functions on the device.

All to say, making the process explicit with something like

```julia
threadCount = Float32(10.) # accessible on the host
threadCountDevice = cu(threadCount) # CuFloat32 accessible on the device

```

is at least an interesting thought to me for the sake of consistency. I’m sure there’s a good reason, I just find the inconsistency interesting.
