# Understanding GPU Kernels

**URL:** <https://discourse.julialang.org/t/understanding-gpu-kernels/10252>\
**Category:** GPU\
**Created:** [April 10, 2018, 12:30am UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252 "2018-04-10T00:30:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jacobcvt12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jacobcvt12/32/2695_2.png) [@jacobcvt12](https://discourse.julialang.org/u/jacobcvt12)\
**Post date:** [April 10, 2018, 12:30am UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252/1 "2018-04-10T00:30:51Z")

</div>

I’m trying understand the basics of GPU kernels for Julia. I’ve followed [http://mikeinnes.github.io/2017/08/24/cudanative.html](http://mikeinnes.github.io/2017/08/24/cudanative.html) and the basic example for addition makes sense

```julia
using CuArrays, CUDAnative

n = 1024
xs, ys, zs = CuArray(rand(n)), CuArray(rand(n)), CuArray(zeros(n))

function kernel_vadd(out, a, b)
  i = (blockIdx().x-1) * blockDim().x + threadIdx().x
  out[i] = a[i] + b[i]
  return
end

@cuda (1, n) kernel_vadd(zs, xs, ys)

```

On my graphics card, I’m limited to 1,024 threads and 4GB of video memory. My question is: if I package up some GPU function for others to use, how can I “automatically” determine how many blocks/threads to allocate on his/her graphics card? Additionally, what is the best practice for “wrapping” this `@cuda(blocks, threads) function...` code?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [April 10, 2018, 6:54am UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252/2 "2018-04-10T06:54:44Z")

</div>

> [@jacobcvt12](#):
>
> My question is: if I package up some GPU function for others to use, how can I “automatically” determine how many blocks/threads to allocate on his/her graphics card?

Insofar your algorithm allows for arbitrary launch parameters, you can query device limits using CUDAdrv. For example, see [the CUDAnative pairwise example](https://github.com/JuliaGPU/CUDAnative.jl/blob/2271487e4fd493c6a66a9b9bb48110fa6f6b6a1e/examples/pairwise.jl#L85-L99):

```julia
total_threads = min(n, attribute(dev, CUDAdrv.MAX_THREADS_PER_BLOCK))

```

However, the best launch configuration depends on more than only the max threads and available memory. Often you want to maximize occupancy, which also depends on the kernel register pressure, cache behavior and/or shared memory usage.

> [@jacobcvt12](#):
>
> Additionally, what is the best practice for “wrapping” this @cuda(blocks, threads) function… code?

Just wrap it in a function? See [the CUDAnative reduce example](https://github.com/JuliaGPU/CUDAnative.jl/blob/2271487e4fd493c6a66a9b9bb48110fa6f6b6a1e/examples/reduce/reduce.jl#L82-L102), although for real-life usage this function should probably also accept a stream parameter.

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [April 10, 2018, 9:17am UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252/3 "2018-04-10T09:17:55Z")

</div>

Or you use `GPUArrays.gpu_call`, which has the added benefit that it can also run with CLArrays.

```julia

# Could also be CLArrays + CLArray
using CuArrays, GPUArrays
n = 1024
xs, ys, zs = CuArray(rand(n)), CuArray(rand(n)), CuArray(zeros(n))
dispatch_dummy = zs # for stream + dispatch to correct backend
args = (zs, xs, ys)
gpu_call(dispatch_dummy, args) do gpu_state, out, a, b
    i = linear_index(gpu_state)
    if i <= length(out)
        out[i] = a[i] + b[i]
    end
    return
end

```

With `gpu_call` the launch parameters are optional:  
[https://juliagpu.github.io/GPUArrays.jl/latest/#GPUArrays.gpu\_call](https://juliagpu.github.io/GPUArrays.jl/latest/#GPUArrays.gpu_call)  
And default to something reasonable, but as @maleadt indicated, you might need manual tuning.

---

<div class="post-metadata">

**Author:** ![jacobcvt12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jacobcvt12/32/2695_2.png) [@jacobcvt12](https://discourse.julialang.org/u/jacobcvt12)\
**Post date:** [April 10, 2018, 1:50pm UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252/4 "2018-04-10T13:50:58Z")

</div>

Thanks @maleadt and @sdanisch! This was just what I was looking for.

---

<div class="post-metadata">

**Author:** ![MikeInnes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikeinnes/32/3656_2.png) [@MikeInnes](https://discourse.julialang.org/u/MikeInnes)\
**Post date:** [April 10, 2018, 3:05pm UTC](https://discourse.julialang.org/t/understanding-gpu-kernels/10252/5 "2018-04-10T15:05:38Z")

</div>

Just to note, there’s nothing magical about the [CuArrays](https://github.com/JuliaGPU/CuArrays.jl) implementation, it’s all just the same constructs – so it might be worth poking through if you’re interested in putting similar things together. If you write something that’s reasonably generic we’d even be happy to take PRs for it.
