# Weird behavior in @cuda kernel dispatch

**URL:** <https://discourse.julialang.org/t/weird-behavior-in-cuda-kernel-dispatch/55724>\
**Category:** GPU\
**Tags:** cuda, kernel\
**Created:** [February 21, 2021, 3:53pm UTC](https://discourse.julialang.org/t/weird-behavior-in-cuda-kernel-dispatch/55724 "2021-02-21T15:53:57Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shuhua](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shuhua/32/27618_2.png) [@Shuhua](https://discourse.julialang.org/u/Shuhua)\
**Post date:** [February 21, 2021, 3:53pm UTC](https://discourse.julialang.org/t/weird-behavior-in-cuda-kernel-dispatch/55724/1 "2021-02-21T15:53:57Z")

</div>

Consider the following dummy kernel function that accepts an argument of type `CuVector`.

```julia
julia> using CUDA

julia> function my_kernel(v::CuVector)
           nothing
       end
my_kernel (generic function with 1 method)

julia> v = CUDA.randn(10)

julia> v isa CuVector
true

julia> my_kernel(v)

```

The above code works well as expected. However, if I launch the kernel with `@cuda`

```julia
julia> @cuda my_kernel(v)
ERROR: MethodError: no method matching my_kernel(::CuDeviceVector{Float32, 1})
Closest candidates are:
  my_kernel(::CuArray{T, 1} where T) at REPL[2]:1
Stacktrace:
 [1] macro expansion
   @ ~/.julia/packages/GPUCompiler/XwWPj/src/cache.jl:0 [inlined]
 [2] cached_compilation(cache::Dict{UInt64, Any}, job::GPUCompiler.CompilerJob{GPUCompiler.PTXCompilerTarget, CUDA.CUDACompilerParams, GPUCompiler.FunctionSpec{typeof(my_kernel), Tuple{CuDeviceVector{Float32, 1}}}}, compiler::typeof(CUDA.cufunction_compile), linker::typeof(CUDA.cufunction_link))
   @ GPUCompiler ~/.julia/packages/GPUCompiler/XwWPj/src/cache.jl:70
 [3] cufunction(f::typeof(my_kernel), tt::Type{Tuple{CuDeviceVector{Float32, 1}}}; name::Nothing, kwargs::Base.Iterators.Pairs{Union{}, Union{}, Tuple{}, NamedTuple{(), Tuple{}}})
   @ CUDA ~/.julia/packages/CUDA/Zmd60/src/compiler/execution.jl:294
 [4] cufunction(f::typeof(my_kernel), tt::Type{Tuple{CuDeviceVector{Float32, 1}}})
   @ CUDA ~/.julia/packages/CUDA/Zmd60/src/compiler/execution.jl:288
 [5] top-level scope
   @ ~/.julia/packages/CUDA/Zmd60/src/compiler/execution.jl:102

```

Why does it complain no matching method? Note that `v isa CuVector` returns `true`.

Verison info:

```julia
Julia Version 1.6.0-rc1
Commit a58bdd9010 (2021-02-06 15:49 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Xeon(R) CPU E5-1603 v3 @ 2.80GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-11.0.1 (ORCJIT, haswell)

```

and CUDA v2.6.1.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [February 21, 2021, 5:57pm UTC](https://discourse.julialang.org/t/weird-behavior-in-cuda-kernel-dispatch/55724/2 "2021-02-21T17:57:51Z")

</div>

`CuVector` and `CuArray` are _host_ representations of cuda memory buffers. Because of this, they carry a lot of metadata and extraneous fields that don’t make any sense in the context of a cuda kernel (which expects little more than an unadorned pointer to some device memory). That’s the role `CuDeviceVector` and `CuDeviceArray` fill.

When you invoke a kernel with `@cuda`, CUDA.jl will auto-convert between host and device arrays. However, since `my_kernel` only accepts host arrays, julia will throw a MethodError when CUDA.jl tries to invoke it as GPU-side code with a device array. The easiest way to resolve this would be to remove the type annotation from `my_kernel`, since you’re likely only ever going to be calling it via `@cuda` anyhow.

---

<div class="post-metadata">

**Author:** ![Shuhua](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shuhua/32/27618_2.png) [@Shuhua](https://discourse.julialang.org/u/Shuhua)\
**Post date:** [February 22, 2021, 2:51am UTC](https://discourse.julialang.org/t/weird-behavior-in-cuda-kernel-dispatch/55724/3 "2021-02-22T02:51:21Z")

</div>

Many thanks for the explanation. You are totally right.

In the documentation of `@cuda` [https://juliagpu.gitlab.io/CUDA.jl/api/compiler/#CUDA.@cuda](https://juliagpu.gitlab.io/CUDA.jl/api/compiler/#CUDA.@cuda) there is

> It will be compiled to a CUDA function upon first use, and to a certain extent arguments will be converted and managed automatically using `cudaconvert` .

while

```julia
julia> CUDA.cudaconvert(CUDA.rand(3)) |> typeof

CuDeviceVector{Float32, 1} (alias for CuDeviceArray{Float32, 1, 1})

```

In addition

```julia
julia> supertype(CuDeviceVector)
AbstractVector{T} where T (alias for AbstractArray{T, 1} where T)

```

An alternative way is thus to declare the argument type to be `AbstractVector`.
