# CUDAnative: Using second and third dims in the kernel

**URL:** <https://discourse.julialang.org/t/cudanative-using-second-and-third-dims-in-the-kernel/1799>\
**Category:** GPU\
**Tags:** cudanative\
**Created:** [January 31, 2017, 1:03pm UTC](https://discourse.julialang.org/t/cudanative-using-second-and-third-dims-in-the-kernel/1799 "2017-01-31T13:03:35Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Colin\_Beckingham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/colin_beckingham/32/4878_2.png) [@Colin\_Beckingham](https://discourse.julialang.org/u/Colin_Beckingham)\
**Post date:** [January 31, 2017, 1:03pm UTC](https://discourse.julialang.org/t/cudanative-using-second-and-third-dims-in-the-kernel/1799/1 "2017-01-31T13:03:35Z")

</div>

In my attempts to generalize the example code given on the README.md I have had good results including processing arrays larger than the device blocksize, but have run into a block regarding the use of the second and third dimensions in the kernel.

I believe the code as presented processes the 3x4 array of floats as a 1x12 vector with the transfer back to an array somehow accomplished in the background. I guess I have a few questions: is it wise to ensure that the thread index does not exceed the dims of the known working area or is that taken care of transparently? My attempts to pass up a tuple for the grid dimensions in @cuda(…) were successful, but I am not getting sensible results from these efforts. I wonder if it might be helpful to include an example where we use at least the y component of the {x,y,z} set?

I can post some code if of interest.

---

<div class="post-metadata">

**Author:** ![Colin\_Beckingham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/colin_beckingham/32/4878_2.png) [@Colin\_Beckingham](https://discourse.julialang.org/u/Colin_Beckingham)\
**Post date:** [January 31, 2017, 1:39pm UTC](https://discourse.julialang.org/t/cudanative-using-second-and-third-dims-in-the-kernel/1799/2 "2017-01-31T13:39:22Z")

</div>

OK I think I have answered my own question and have it working, thanks.  
However the puzzling thing is that in the kernel I can write it two ways, with the x dim alone or the x and y and both produce a result which passes the test. I’ll get it eventually.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 31, 2017, 3:58pm UTC](https://discourse.julialang.org/t/cudanative-using-second-and-third-dims-in-the-kernel/1799/3 "2017-01-31T15:58:38Z")

</div>

> [@Colin\_Beckingham](#):
>
> is it wise to ensure that the thread index does not exceed the dims of the known working area

Depends on the kernel preconditions. If you know that the index calculation will never yield an out-of-bounds index, you don’t have to. But when generalizing for larger arrays, that might not be possible (ie. 513 items on a max-512-threads device == 512 threads 2 blocks).

There’s a lot of existing literature on how to flexibly generalize kernels, e.g. [writing grid-stride loops](https://devblogs.nvidia.com/parallelforall/cuda-pro-tip-write-flexible-kernels-grid-stride-loops/). You should also [take care](http://stackoverflow.com/questions/9985912/how-do-i-choose-grid-and-block-dimensions-for-cuda-kernels) whether to launch more threads or more blocks, it determines occupancy and consequently performance, but also depends on the kernel and the hardware.

> [@Colin\_Beckingham](#):
>
> in the kernel I can write it two ways, with the x dim alone or the x and y and both produce a result which passes the test

You can use `@cuprintf` to debug your index calculations, see [this example](https://github.com/JuliaGPU/CUDAnative.jl/blob/master/examples/hello_world.jl).
