# Delays shown in Nsight Systems between HtoD memcopy and kernel launch when using CUDA.jl

**URL:** <https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446>\
**Category:** GPU\
**Tags:** gpu, cudajl\
**Created:** [July 24, 2024, 9:45pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446 "2024-07-24T21:45:28Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Cibin\_Joseph](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cibin_joseph/32/208886_2.png) [@Cibin\_Joseph](https://discourse.julialang.org/u/Cibin_Joseph)\
**Post date:** [July 24, 2024, 9:45pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/1 "2024-07-24T21:45:28Z")

</div>

For the CUDA.jl minimal example shown below, I see a considerably large delay between the HtoD memcopy and the kernel launch. I’m not able to figure out what causes it. For a larger array, _sometimes_ this delay is much lesser too. Is this part of the CUDA API calls, and something to be expected?

```julia
using CUDA
using Random

@inline function kernel!(a)
    i = threadIdx().x + blockDim().x*(blockIdx().x-1)
    a[i] = CUDA.cos(a[i]) + i^0.6 + CUDA.tan(a[i])
    return
end

function main()
    n = 2^12
    a = rand(n)
    
    threads = 256
    blocks = cld(n, threads)
    a_d = CuArray(a)
    @cuda threads=threads blocks=blocks kernel!(a_d)
    a = Array(a_d)

    return
end

main()
CUDA.@profile main()

```

The times reported by Nsight systems are:  
HtoD memcopy: ~3 microseconds  
**Delay: ~157 microseconds**  
Kernel: ~18 microseconds  
DtoH memcopy: ~3 microseconds

 ![Screenshot from 2024-07-24 15-30-45](https://global.discourse-cdn.com/julialang/original/3X/3/9/39a20d8b74236d01e1bb823e2a6f205eea10ace2.png)

Additionally, for debugging scenarios like these, is there a way to make Nsight Systems record the CPU operations and function calls too?

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [July 24, 2024, 10:44pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/2 "2024-07-24T22:44:17Z")

</div>

> [@Cibin\_Joseph](#):
>
> is there a way to make Nsight Systems record the CPU operations and function calls too?

> **[GitHub - JuliaGPU/NVTX.jl: Julia bindings for NVTX, for instrumenting with...](https://github.com/JuliaGPU/NVTX.jl)**
>
> Julia bindings for NVTX, for instrumenting with the Nvidia Nsight Systems profiler - JuliaGPU/NVTX.jl

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [July 25, 2024, 6:12pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/3 "2024-07-25T18:12:28Z")

</div>

In addition, this makes it possible to figure out whether this is a GC stall ([NVTX.jl · NVTX.jl](https://juliagpu.github.io/NVTX.jl/dev/#Instrumenting-Julia-internals)), although I don’t see any free operations being queued in that gap, so it seems unlikely.

---

<div class="post-metadata">

**Author:** ![Cibin\_Joseph](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cibin_joseph/32/208886_2.png) [@Cibin\_Joseph](https://discourse.julialang.org/u/Cibin_Joseph)\
**Post date:** [July 29, 2024, 3:14am UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/4 "2024-07-29T03:14:17Z")

</div>

The NVTX annotations were helpful in determining where exactly the kernel launch begins and when the HtoD ends. I also did not get any blocks that had GC in the Nsight profiles. This is a fairly small case, so I’m ruling that out for now.  
I still find that there’s a substantial delay between memcopy (both HtoD and DtoH) and kernel execution. I’m tempted to think there’s something the Julia API is stuck with.

If I wanted to try to figure this out? Do you guys have any suggestions on how I could go about doing that? Would it be worth it looking into the llvm code?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [July 29, 2024, 9:34am UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/5 "2024-07-29T09:34:43Z")

</div>

I would recommend adding `NVTX.@annotate` annotations to functions, both in your application and in CUDA.jl (after `]dev`ing the package)… In combination with interactive profiling and Revise.jl this should make it possible to quickly find where the delay is coming from.

---

<div class="post-metadata">

**Author:** ![Cibin\_Joseph](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cibin_joseph/32/208886_2.png) [@Cibin\_Joseph](https://discourse.julialang.org/u/Cibin_Joseph)\
**Post date:** [July 29, 2024, 10:33pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/6 "2024-07-29T22:33:42Z")

</div>

I tried going through and annotating the cufunction to see why it was taking so long to launch the GPU kernel and it seems like the call to `methodinstance` and `cached_compilation` takes quite some time. At this point, I’m realizing I’m way too deep down this rabbit hole. So, I’m going to sketch it up to compilation overheads and move on. Thanks!

```julia
function cufunction(f::F, tt::TT=Tuple{}; kwargs...) where {F,TT}
    cuda = active_state()

    Base.@lock cufunction_lock begin
        # compile the function
        cache = compiler_cache(cuda.context)
        NVTX.@mark "compiler_cache"
        source = methodinstance(F, tt)
        NVTX.@mark "methodinstance"
        config = compiler_config(cuda.device; kwargs...)::CUDACompilerConfig
        NVTX.@mark "compiler_config"
        fun = GPUCompiler.cached_compilation(cache, source, config, compile, link)
        
        NVTX.@mark "cached_compilation"
        # create a callable object that captures the function instance. we don't need t
o think 
        # about world age here, as GPUCompiler already does and will return a different
 object 
        key = (objectid(source), hash(fun), f)
        NVTX.@mark "hash"
        kernel = get(_kernel_instances, key, nothing)
        NVTX.@mark "get"
        if kernel === nothing
            # create the kernel state object
            state = KernelState(create_exceptions!(fun.mod), UInt32(0))
            NVTX.@mark "kernelstate"
            
            kernel = HostKernel{F,tt}(f, fun, state)
            NVTX.@mark "hostkernel"
            _kernel_instances[key] = kernel
        end 
        NVTX.@mark "after if condition"
        return kernel::HostKernel{F,tt}
    end 
end 

```

 ![Screenshot from 2024-07-29 15-05-46](https://global.discourse-cdn.com/julialang/original/3X/2/2/2236b85a72bf9615cf73ee33ea85d21c43d1acf4.png)

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [July 31, 2024, 11:31am UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/7 "2024-07-31T11:31:37Z")

</div>

> [@Cibin\_Joseph](#):
>
> it seems like the call to `methodinstance` and `cached_compilation` takes quite some time

Hmm, that shouldn’t be happening. Which version of Julia are you using? On 1.10+, those lookups should be fast.

---

<div class="post-metadata">

**Author:** ![Cibin\_Joseph](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cibin_joseph/32/208886_2.png) [@Cibin\_Joseph](https://discourse.julialang.org/u/Cibin_Joseph)\
**Post date:** [July 31, 2024, 3:14pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/8 "2024-07-31T15:14:21Z")

</div>

I’m working on Julia 1.10.4

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [July 31, 2024, 8:27pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/9 "2024-07-31T20:27:29Z")

</div>

In that case, can you file an issue on CUDA.jl with the reproducer from the top, your findings on which functions behave badly, and what versions of packages you were using exactly (i.e., a Manifest)?

---

<div class="post-metadata">

**Author:** ![Cibin\_Joseph](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cibin_joseph/32/208886_2.png) [@Cibin\_Joseph](https://discourse.julialang.org/u/Cibin_Joseph)\
**Post date:** [July 31, 2024, 11:22pm UTC](https://discourse.julialang.org/t/delays-shown-in-nsight-systems-between-htod-memcopy-and-kernel-launch-when-using-cuda-jl/117446/10 "2024-07-31T23:22:04Z")

</div>

Github issue [#2456](https://github.com/JuliaGPU/CUDA.jl/issues/2456) opened in CUDA.jl
