# Performance of kernel function

**URL:** <https://discourse.julialang.org/t/performance-of-kernel-function/31605>\
**Category:** GPU\
**Created:** [November 28, 2019, 5:00am UTC](https://discourse.julialang.org/t/performance-of-kernel-function/31605 "2019-11-28T05:00:37Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ennvvy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ennvvy/32/7816_2.png) [@ennvvy](https://discourse.julialang.org/u/ennvvy)\
**Post date:** [November 28, 2019, 5:00am UTC](https://discourse.julialang.org/t/performance-of-kernel-function/31605/1 "2019-11-28T05:00:37Z")

</div>

I tried executing the following function on the GPU, for array sizes of 10,000.

```julia
function update!(s,c,cedg,θ)
  index = (blockIdx().x - 1) * blockDim().x + threadIdx().x
  stride = blockDim().x * gridDim().x
  @inbounds for l=index:stride:length(s)
                s[l]=cedg[l]*CUDAnative.sin(θ[l])
                c[l]=cedg[l]*CUDAnative.cos(θ[l])
            end
end   

```

However, the performance is similar to the code run on the CPU. Is there something wrong in the way it is written?

---

<div class="post-metadata">

**Author:** ![ennvvy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ennvvy/32/7816_2.png) [@ennvvy](https://discourse.julialang.org/u/ennvvy)\
**Post date:** [November 28, 2019, 5:11am UTC](https://discourse.julialang.org/t/performance-of-kernel-function/31605/3 "2019-11-28T05:11:43Z")

</div>

Sorry, I had posted a different version. I have now updated it with the one I use.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [November 28, 2019, 5:19am UTC](https://discourse.julialang.org/t/performance-of-kernel-function/31605/4 "2019-11-28T05:19:26Z")

</div>

```julia
s = zeros(10000);
c = zeros(10000);
cedg = rand((0,1),10000) .* randn(10000);
θ = randn(10000);
cs = cu(s);
cc = cu(c);
ccedg = cu(cedg);
cθ = cu(θ);

function update!(s,c,cedg,θ)
  @inbounds for l=eachindex(s)
    s[l]=cedg[l]*sin(θ[l])
    c[l]=cedg[l]*cos(θ[l])
  end
end
function update!(s::CuArray,c,cedg,θ)
  s .= cedg.*sin.(θ)
  c .= cedg.*cos.(θ)
end

@btime update!($s,$c,$cedg,$θ);
@btime update!($cs,$cc,$ccedg,$cθ);

julia> @btime update!($s,$c,$cedg,$θ);

  132.067 μs (0 allocations: 0 bytes)

julia> @btime update!($cs,$cc,$ccedg,$cθ);
  10.649 μs (108 allocations: 4.38 KiB)

```

The vectors have to be long enough for it to be worth it though

---

<div class="post-metadata">

**Author:** ![ennvvy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ennvvy/32/7816_2.png) [@ennvvy](https://discourse.julialang.org/u/ennvvy)\
**Post date:** [November 28, 2019, 5:35am UTC](https://discourse.julialang.org/t/performance-of-kernel-function/31605/5 "2019-11-28T05:35:11Z")

</div>

@baggepinnen: Thank you. This worked, I was trying to follow the tutorial on GPU programming using CuArrays and was trying to fit my functions like the ones in the example. One clarification, do I not have to specify the number of threads and blocks or even mention `@cuda` for it to be executed on the GPU?
