# KernelAbstractions Autotuning

**URL:** <https://discourse.julialang.org/t/kernelabstractions-autotuning/108013>\
**Category:** GPU\
**Created:** [December 24, 2023, 6:48pm UTC](https://discourse.julialang.org/t/kernelabstractions-autotuning/108013 "2023-12-24T18:48:35Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![smartalecH](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smartalech/32/32379_2.png) [@smartalecH](https://discourse.julialang.org/u/smartalecH)\
**Post date:** [December 24, 2023, 6:48pm UTC](https://discourse.julialang.org/t/kernelabstractions-autotuning/108013/1 "2023-12-24T18:48:35Z")

</div>

I’m going through the recent [paper](https://arxiv.org/pdf/2304.06835.pdf) comparing batched KernelAbstractions kernels to standard array batching in jax and pytorch. In section 5.1.2, I found this interesting tidbit:

> KernelAbstractions.jl performs a limited form of auto-tuning by optimizing the launch parameters for occupancy.

I went back to the docs to see if I could find anything describing this, but came up empty handed.

Perhaps I’m looking in the wrong place. Does anyone have any references that describe this functionality (and how well it works across different hardware platforms)?

Thanks!

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 25, 2023, 11:48am UTC](https://discourse.julialang.org/t/kernelabstractions-autotuning/108013/2 "2023-12-25T11:48:29Z")

</div>

It is backend dependent, but if you don’t specify the workgroupsize the back ends makes an educated guess.

As an example CUDA does [CUDA.jl/src/CUDAKernels.jl at 3605167a9ea3aebfc944cc88ea0f86f01723a764 · JuliaGPU/CUDA.jl · GitHub](https://github.com/JuliaGPU/CUDA.jl/blob/3605167a9ea3aebfc944cc88ea0f86f01723a764/src/CUDAKernels.jl#L123)
