# Is it possible to index a CuArray with a CuArray?

**URL:** <https://discourse.julialang.org/t/is-it-possible-to-index-a-cuarray-with-a-cuarray/18812>\
**Category:** GPU\
**Tags:** question\
**Created:** [December 18, 2018, 9:53pm UTC](https://discourse.julialang.org/t/is-it-possible-to-index-a-cuarray-with-a-cuarray/18812 "2018-12-18T21:53:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Roger-luo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roger-luo/32/3399_2.png) [@Roger-luo](https://discourse.julialang.org/u/Roger-luo)\
**Post date:** [December 18, 2018, 9:53pm UTC](https://discourse.julialang.org/t/is-it-possible-to-index-a-cuarray-with-a-cuarray/18812/1 "2018-12-18T21:53:39Z")

</div>

The case is that I have a data set of indexes (rather than vectors), so I need to frequently index a multi-dimensional array (which is the table), e.g

```julia
A = [rand(3, 3, 2) for i in 1:8]
data = rand(1:2, 8, 10) # last dimension is the length of data set

my_task(A, data) = [A[k][:, :, data[k, :]] for k in 1:8]

```

this (batched indexing) works with `Array`, but not with `CuArray`, I’m wondering if I could parallel the indexing on GPU without copying it back to CPU (since the batch size of indexes would be pretty large, and easy to parallel).

---

<div class="post-metadata">

**Author:** ![1115](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/1115/32/4465_2.png) [@1115](https://discourse.julialang.org/u/1115)\
**Post date:** [January 11, 2019, 5:30am UTC](https://discourse.julialang.org/t/is-it-possible-to-index-a-cuarray-with-a-cuarray/18812/2 "2019-01-11T05:30:09Z")

</div>

You can write a simple implementation with CUDAnative.jl, e.g.

```julia
function getindex(A::CuVector{T}, B::CuArray{<:Integer}) where T
    res = cuzeros(T, size(B)...)
    @inline function kernel(res, A, B)
        state = (blockIdx().x-1) * blockDim().x + threadIdx().x
        state <= length(res) && (@inbounds res[state] = A[B[state]])
        return
    end

    max_threads = 256
    X, Y = thread_blocks_heuristic(length(B))
    @cuda threads=X blocks=Y kernel(res, A, B)
    res
end

```
