# Any qualifier in CUDA.jl like \`\_\_device\_\_\` in CUDA/C++?

**URL:** <https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517>\
**Category:** GPU\
**Tags:** question\
**Created:** [April 25, 2024, 11:12pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517 "2024-04-25T23:12:55Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![huiyuxie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/huiyuxie/32/214265_2.png) [@huiyuxie](https://discourse.julialang.org/u/huiyuxie)\
**Post date:** [April 25, 2024, 11:12pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/1 "2024-04-25T23:12:55Z")

</div>

Hi all! I was wondering is there any qualifier in CUDA.jl like ` __device__ ` in CUDA/C++ to calling a function from device to make sure the function call is not falling back to CPU.

For example,

```julia
using CUDA

function gpu_add!(a, b, c)
    i = threadIdx().x
    a[i] = func(b[i], c[i])
    return
end

```

`func` is originally defined on CPU, will launching this `gpu_add!` kernel fall back to CPU when calling `func`? If it is going to fall back, is there any good macro equivalent to ` __device__ ` in CUDA/C++ to specifically restrict a function being running solely on GPU. Thanks!

Due to some reasons, there are so many cases like above in my project, but they are all numerical function (i.e., no way to parallelize them using parallel computing). If falling back is unavoidable, will it significantly affect performance? Thanks again!

---

<div class="post-metadata">

**Author:** ![Zentrik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zentrik/32/35409_2.png) [@Zentrik](https://discourse.julialang.org/u/Zentrik)\
**Post date:** [April 25, 2024, 11:24pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/2 "2024-04-25T23:24:12Z")

</div>

As far as I’m aware, in a kernel everything executes on the GPU.  
I believe CUDA.jl has an internal macro `@device_function` that is the analogue of ` __device__ `.

---

<div class="post-metadata">

**Author:** ![huiyuxie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/huiyuxie/32/214265_2.png) [@huiyuxie](https://discourse.julialang.org/u/huiyuxie)\
**Post date:** [April 25, 2024, 11:33pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/3 "2024-04-25T23:33:10Z")

</div>

Thanks! So as a user of CUDA.jl, can we assume that everything runs on the GPU within a kernel without worrying about falling back to the CPU in accident?

---

<div class="post-metadata">

**Author:** ![SteffenPL](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steffenpl/32/206270_2.png) [@SteffenPL](https://discourse.julialang.org/u/SteffenPL)\
**Post date:** [April 26, 2024, 1:28am UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/4 "2024-04-26T01:28:12Z")

</div>

I believe as long as you call a function with `@cuda ...` it will be fully compiled by the GPUCompiler.jl and hence any operations which would be infeasible on the GPU will lead directly to a compiler error. This includes function calls inside that function, which all will be compiled on the GPU as well.

(Note 100% sure, but maybe it is better to think of a function definition as something independent of the CPU or GPU, as Julia can compile different versions (methods) for a function depending on the inputs when the function is called. Therefore, it might happen that the function `func` will never be compiler for the CPU if it is only used inside the GPU kernel.)

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [April 26, 2024, 1:48am UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/5 "2024-04-26T01:48:04Z")

</div>

Yes, everything called within a kernel will be executed on the GPU and if we are unable to compile it we will throw an error.

---

<div class="post-metadata">

**Author:** ![huiyuxie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/huiyuxie/32/214265_2.png) [@huiyuxie](https://discourse.julialang.org/u/huiyuxie)\
**Post date:** [April 26, 2024, 4:19pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/6 "2024-04-26T16:19:19Z")

</div>

Thanks!

---

<div class="post-metadata">

**Author:** ![huiyuxie](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/huiyuxie/32/214265_2.png) [@huiyuxie](https://discourse.julialang.org/u/huiyuxie)\
**Post date:** [April 26, 2024, 4:20pm UTC](https://discourse.julialang.org/t/any-qualifier-in-cuda-jl-like-device-in-cuda-c/113517/7 "2024-04-26T16:20:51Z")

</div>

Thanks for the reply!
