# Column-wise reduction on a CUDA.CuArray matrix

**URL:** <https://discourse.julialang.org/t/column-wise-reduction-on-a-cuda-cuarray-matrix/45069>\
**Category:** Performance\
**Tags:** question, cudanative, indexing, cuda, cuarrays\
**Created:** [August 16, 2020, 10:33pm UTC](https://discourse.julialang.org/t/column-wise-reduction-on-a-cuda-cuarray-matrix/45069 "2020-08-16T22:33:44Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkarikom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkarikom/32/7096_2.png) [@mkarikom](https://discourse.julialang.org/u/mkarikom)\
**Post date:** [August 16, 2020, 10:33pm UTC](https://discourse.julialang.org/t/column-wise-reduction-on-a-cuda-cuarray-matrix/45069/1 "2020-08-16T22:33:44Z")

</div>

As suggested [here](https://discourse.julialang.org/t/cuarray-find-first-negative-value-along-columns/45012/2), I am trying to create a version of [this findfirst kernel](https://github.com/JuliaGPU/CUDA.jl/blob/f2ee969226759fc15c8092dcd180f155f7305269/src/indexing.jl#L101) that operates over dimension 2 of input matrix `xs` and returns the first match for each vector along dimension 1.

I think a good approach is [a column-wise reduction as discussed here](https://forums.developer.nvidia.com/t/get-first-row-index-that-meets-the-condition-in-cuda/42316).  
Note that this is ported from a suggestion offered by one of the mods at the nvidia dev forum, who claims that it launches one thread per column, which I don’t

As a first step, I tried finding the first row of each column of `arr` \> 0.5, with `size(arr)` constant:

```julia
function kernel(arr::CuDeviceArray,result::CuDeviceArray)
    thresh = Float32(0.5)
    nrow = Int32(3)
    ncol = Int32(100)
    idx = threadIdx().x+blocDim().x*blocIdx().x
    if idx < ncol
      res = Int32(-1)
      i = int32(0)
      while i < nrow && res == Int32(-1)
        if arr[i] > thresh
          res = i
        end
        result[idx] = res
      end
    end
end

```

Now launch this setting `nthreads` to the number of rows and `nblocks` to the number of columns, so that all the columns get processed in parallel.  
Get a weird error:

```julia

check = CUDA.rand(3,100)
nthreads = size(check)[1]
nblocks = size(check)[2]
rslt = CUDA.fill(Int32(0),size(check)[2])
@cuda threads=nthreads blocks=nblocks kernel(check,rslt)

ERROR: MethodError: no method matching kernel(::CuDeviceArray{Float32,2,CUDA.AS.Global}, ::CuDeviceArray{Int32,1,CUDA.AS.Global})

```

Obviously the `CuArray` arguments to `kernel()` are getting passed as `CuDeviceArray` with `ptr::CUDA.AS.Global` and `CUDA.jl` wanted some other value for the `ptr` parameter.

I’m now stuck on what the `ptr` parameter means and how `CuArray` becomes `CuDeviceArray`.
