# Dynamic parallelism slow in CUDA.jl

**URL:** <https://discourse.julialang.org/t/dynamic-parallelism-slow-in-cuda-jl/117463>\
**Category:** GPU\
**Created:** [July 25, 2024, 10:39am UTC](https://discourse.julialang.org/t/dynamic-parallelism-slow-in-cuda-jl/117463 "2024-07-25T10:39:53Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![z-wang](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/z-wang/32/210507_2.png) [@z-wang](https://discourse.julialang.org/u/z-wang)\
**Post date:** [July 25, 2024, 10:39am UTC](https://discourse.julialang.org/t/dynamic-parallelism-slow-in-cuda-jl/117463/1 "2024-07-25T10:39:53Z")

</div>

Hi, i’m learning dynamic parallelism in `CUDA.jl` . Here i have a code which executes a parent kernel concurrently 1000 times, and each parent kernel queues the child kernel `N` times. My problem is that this code is pretty slow, and the running time increases exponentially with `N`, even though the child kernel literally does nothing.

Could someone help me where I was doing wrong here?

```julia
using CUDA

function example_parent(N)
    for i in 1:N
        @cuda threads = 100 dynamic = true example_child()
    end
    return nothing
end

function example_child()
    return nothing
end

function test(N)
    CUDA.@sync begin
        @cuda threads = 1000 example_parent(N)
    end
end

test(2) # run once to compile
CUDA.@time test(2) # 0.052070 seconds (9 CPU allocations: 416 bytes)
CUDA.@time test(3) # 6.671022 seconds (397 CPU allocations: 25.188 KiB)

```

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [July 25, 2024, 6:09pm UTC](https://discourse.julialang.org/t/dynamic-parallelism-slow-in-cuda-jl/117463/2 "2024-07-25T18:09:50Z")

</div>

> **[CUDA Dynamic Parallelism API and Principles | NVIDIA Technical Blog](https://developer.nvidia.com/blog/cuda-dynamic-parallelism-api-principles/)**
>
> This post is the second in a series on CUDA Dynamic Parallelism. In my first post, I introduced Dynamic Parallelism by using it to compute images of the Mandelbrot set using recursive subdivision…

> By default, space is reserved for 2048 pending child grids; this can be extended by setting the appropriate device limit, as in the following code.  
> …  
> The runtime first tries to add the newly launched grid to the fixed-size pool, and if it is full, uses the virtualized pool. While this means that grids are queued successfully, the costs of using the virtualized pool are higher than those of the fixed-size pool.
