# Optimize code which uses KernelAbstractions.jl

**URL:** <https://discourse.julialang.org/t/optimize-code-which-uses-kernelabstractions-jl/110068>\
**Category:** Performance\
**Tags:** gpu, kernelabstractions\
**Created:** [February 11, 2024, 4:24pm UTC](https://discourse.julialang.org/t/optimize-code-which-uses-kernelabstractions-jl/110068 "2024-02-11T16:24:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [February 11, 2024, 4:24pm UTC](https://discourse.julialang.org/t/optimize-code-which-uses-kernelabstractions-jl/110068/1 "2024-02-11T16:24:00Z")

</div>

Hi,

in [RadonKA.jl](https://github.com/roflmaostc/RadonKA.jl/blob/c266241d7c3d3726e510f8f6aa2c08b7d9f56f1b/src/radon.jl#L171) I just launch the kernel with the following code:

```julia
[....]
    kernel! = radon_kernel!(backend)
    kernel!(sinogram::AbstractArray{T}, img, weights, in_height, 
            out_height, angles, mid, radius, absorb_f,
            ndrange=size(sinogram))
    KernelAbstractions.synchronize(backend)    
    return sinogram::typeof(img)
end

@kernel function radon_kernel!(sinogram::AbstractArray{T}, img::AbstractArray{T}, 
                               weights, in_height, out_height, angles, mid,
                               radius, absorb_f) where {T}
    i, iangle, i_z = @index(Global, NTuple)

```

I was wondering, because in the [KA docs](https://juliagpu.github.io/KernelAbstractions.jl/stable/examples/naive_transpose/) the `groupsize` is mentioned.  
Should I care about it? And which reasonable value do I choose?  
My arrays range from sizes like `(256,256)` to 3D arrays such as `(512,512,512)`.

I also tried annotating `@Const` all arguments except `sinogram`. Didn’t improve performance.  
Is there any other free performance tricks I can use?

Thanks!

Felix

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [February 11, 2024, 5:29pm UTC](https://discourse.julialang.org/t/optimize-code-which-uses-kernelabstractions-jl/110068/2 "2024-02-11T17:29:57Z")

</div>

KA uses a limited form of auto-tuning to select the group size. I would recommend the native performance tools from CUDA to look at kernel performance.

> **[Benchmarking & profiling · CUDA.jl](https://cuda.juliagpu.org/stable/development/profiling/)**
>
> Documentation for CUDA.jl.
