# Question about CUDA kernels

**URL:** <https://discourse.julialang.org/t/question-about-cuda-kernels/94345>\
**Category:** GPU\
**Tags:** question\
**Created:** [February 9, 2023, 12:55pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345 "2023-02-09T12:55:19Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![EduMolinero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/edumolinero/32/36991_2.png) [@EduMolinero](https://discourse.julialang.org/u/EduMolinero)\
**Post date:** [February 9, 2023, 12:55pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345/1 "2023-02-09T12:55:19Z")

</div>

Hi!

I’m struggling to implement the following CUDA kernel:

```julia
using CUDA
using LinearAlgebra

function kernel!(A, B, C)
    X = CUDA.size(A)[3]
    Y = CUDA.size(B)[3]
    for j in 1:X
        for i in 1:Y
            @inbounds C[:,:,j] += B[:,:,i] * A[:,:,j]
        end
    end
    return 
end

function benchmark_gpu(A, B, C)
    CUDA.@sync begin
        @cuda threads=256 kernel!(A, B, C)
    end
    return nothing
end

A = CUDA.rand(ComplexF64,4,4,1000)
B = CUDA.rand(ComplexF64,4,4,10)
C = CUDA.zeros(ComplexF64,4,4,1000)

benchmark_gpu(A, B, C)

```

When I run it as a script, it throws the following error:

```julia
ERROR: LoadError: InvalidIRError: compiling kernel #kernel!(CuDeviceArray{ComplexF64, 3, 1}, CuDeviceArray{ComplexF64, 3, 1}, CuDeviceArray{ComplexF64, 3, 1}) resulted in invalid LLVM IR
Reason: unsupported call through a literal pointer (call to ijl_rethrow)
Stacktrace:
 [1] rethrow
   @ ./error.jl:61
 [2] task_local_storage
   @ ./task.jl:294
 [3] allowscalar
   @ ~/.julia/packages/GPUArraysCore/B3xv7/src/GPUArraysCore.jl:58
 [4] kernel!
   @ ~/cuda_tests/discourse.jl:7
....

```

At first I thought that the error was cause by broadcasting during the matrix multiplication step. However, I can do the operation `C[:,:,j] += B[:,:,i] * A[:,:,j]` when I’m inside a Julia REPL session.  
How’s that possible?

I’m aware that scalar indexing is not the best practice when it comes to maximize performance on the GPU, but I would like to understand why the code is failing anyway.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [February 9, 2023, 3:55pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345/2 "2023-02-09T15:55:42Z")

</div>

> [@EduMolinero](#):
>
> At first I thought that the error was cause by broadcasting during the matrix multiplication step. However, I can do the operation `C[:,:,j] += B[:,:,i] * A[:,:,j]` when I’m inside a Julia REPL session.  
> How’s that possible?

Code in your REPL executes in a CPU environment (even though it may in turn call into the GPU), whereas kernel code executes directly on the GPU. You can only perform much more simple operations there, and not the matrix multiplication or broadcast you’re doing.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [February 9, 2023, 4:31pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345/3 "2023-02-09T16:31:47Z")

</div>

If this operation is what you want, it’s easy to do without writing kernels. It is `C[x,y,j] = B[x,z,i] * A[z,y,j]` summed on `z,i,j`, and the sum on `i` can be done first:

```julia
using TensorCore # reshape + one matrix multiplication
boxdot!(C, dropdims(sum(B, dims=3), dims=3), A)

using NNlib # CUDA's gemm_strided_batched!
batched_mul!(C, sum(B, dims=3), A)

```

---

<div class="post-metadata">

**Author:** ![EduMolinero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/edumolinero/32/36991_2.png) [@EduMolinero](https://discourse.julialang.org/u/EduMolinero)\
**Post date:** [February 10, 2023, 3:24pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345/4 "2023-02-10T15:24:00Z")

</div>

Oh I get it. Thank you for the answer!!

---

<div class="post-metadata">

**Author:** ![EduMolinero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/edumolinero/32/36991_2.png) [@EduMolinero](https://discourse.julialang.org/u/EduMolinero)\
**Post date:** [February 10, 2023, 3:33pm UTC](https://discourse.julialang.org/t/question-about-cuda-kernels/94345/5 "2023-02-10T15:33:02Z")

</div>

That does exactly what I wanted.  
Thank you very much!!

However, in order to properly use the `NNlib` option, I needed to load the CUDA version of the library `NNlibCUDA`.
