# Use of CartesianIndices with CUDA?

**URL:** <https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160>\
**Category:** GPU\
**Tags:** indexing\
**Created:** [June 27, 2020, 5:26pm UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160 "2020-06-27T17:26:24Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ferran\_Mazzanti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ferran_mazzanti/32/7458_2.png) [@Ferran\_Mazzanti](https://discourse.julialang.org/u/Ferran_Mazzanti)\
**Post date:** [June 27, 2020, 5:26pm UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/1 "2020-06-27T17:26:24Z")

</div>

Hi,  
I have a function that iterates over all elements of a four-dimenasional array full of CartesianIndices elements

```julia
21×31×41×51 Array{CartesianIndex{4},4}:
[:, :, 1, 1] =
 CartesianIndex(1, 1, 1, 1) … CartesianIndex(1, 31, 1, 1)
 CartesianIndex(2, 1, 1, 1) CartesianIndex(2, 31, 1, 1)
 CartesianIndex(3, 1, 1, 1) CartesianIndex(3, 31, 1, 1)
 CartesianIndex(4, 1, 1, 1) CartesianIndex(4, 31, 1, 1)
 CartesianIndex(5, 1, 1, 1) CartesianIndex(5, 31, 1, 1)
...

```

the fact is the problem itself the function does is 100% parallelizable, as it has to do some independent operations for each element of the array, and do the product at then (not really the product, but similar stuff). That seems to be perfect for CUDA as there are many elements to process, each one independent of the rest.

Now the question is: can I directly work with these CartesianIndex elements in CUDA? Is this implemented? Or should I convert that to a 4-dimensional array amd work with that? In case I shall convert to an array, how do you properly do a nested for loop (one for each dimension of the CartesianIndex, (and here I have 4), taking full advantage of the CUDA parallel capabilities?

Thanks a lot,

Ferran.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [June 27, 2020, 9:33pm UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/2 "2020-06-27T21:33:39Z")

</div>

You can use `CartesianIndex` objects on the GPU – did you try it? But the real answer to your question depends on how you’re going to use them.

EDIT: Which you will need to provide more information about to be able to give a useful anwer.

---

<div class="post-metadata">

**Author:** ![Ferran\_Mazzanti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ferran_mazzanti/32/7458_2.png) [@Ferran\_Mazzanti](https://discourse.julialang.org/u/Ferran_Mazzanti)\
**Post date:** [June 28, 2020, 7:36am UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/3 "2020-06-28T07:36:15Z")

</div>

for instance, this throws an error on my machine

```julia
using CuArrays
using CuArrays.CURAND
using CUDAnative
using CUDAdrv

# include the path to user-defined modules
# 
#push!(LOAD_PATH,homedir()*"/Julia_1/Modules");
#push!(LOAD_PATH,homedir()*"/Julia_1/Modules/RBM");
#push!(LOAD_PATH,homedir()*"/Julia_1/Modules/CUDA");

A = rand(3,4,5)

aux_CI = CartesianIndices(A)

display(aux_CI)
println()
println(size(aux_CI))
println(length(aux_CI))

tot = CuArrays.zeros(length(aux_CI))

function bucle_1(y,tot)
    index = (blockIdx().x - 1) * blockDim().x + threadIdx().x
    stride = blockDim().x * gridDim().x
    for i = index:stride:length(y)
        tot[i] = y[1]+y[2]
    end
end;

numblocks = 256
@cuda threads = 256 blocks = numblocks bucle_1(aux_CI,tot)

```

I typed it fast so it may be me making mistakes, but anyway…

Best regards,

Ferran.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [June 28, 2020, 5:46pm UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/4 "2020-06-28T17:46:11Z")

</div>

`aux_CI` is an array that lives on the CPU. You need to make a CuArray version.

---

<div class="post-metadata">

**Author:** ![Ellipse0934](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ellipse0934/32/4012_2.png) [@Ellipse0934](https://discourse.julialang.org/u/Ellipse0934)\
**Post date:** [July 1, 2020, 7:05am UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/5 "2020-07-01T07:05:25Z")

</div>

Change

```julia
tot[i] = y[1]+y[2]

```

to

```julia
tot[i] = y[i][1]+y[i][2]

```

You can pass `CartesianIndices(x)` directly to GPU kernels, it will pass `OneTo(len)` instead of constructing an array.

---

<div class="post-metadata">

**Author:** ![duanestorti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/duanestorti/32/37374_2.png) [@duanestorti](https://discourse.julialang.org/u/duanestorti)\
**Post date:** [September 5, 2020, 1:24am UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/6 "2020-09-05T01:24:20Z")

</div>

Is there a place to find examples using CartesianIndex objects on GPU? Is it as straightforward as defining the CartesianIndices object, passing it to the kernel function, and then accessing elements using a linear index constructed from the block and thread index values?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [September 7, 2020, 3:01pm UTC](https://discourse.julialang.org/t/use-of-cartesianindices-with-cuda/42160/7 "2020-09-07T15:01:10Z")

</div>

Yes, there’s several kernels like that in CUDA.jl and GPUArrays.jl, e.g., [https://github.com/JuliaGPU/GPUArrays.jl/blob/b988cdcc81011ded7223f250d127a1e544ea2d2a/src/host/broadcast.jl#L53-L72](https://github.com/JuliaGPU/GPUArrays.jl/blob/b988cdcc81011ded7223f250d127a1e544ea2d2a/src/host/broadcast.jl#L53-L72)  
(where `@cartesianidx` is a simple macro that creates an iterator based on the current block/thread index).
