# Improving GPU performance for symbolic regression

**URL:** <https://discourse.julialang.org/t/improving-gpu-performance-for-symbolic-regression/108800>\
**Category:** Performance\
**Tags:** question, gpu, cuda, symbolic-regression\
**Created:** [January 14, 2024, 9:29pm UTC](https://discourse.julialang.org/t/improving-gpu-performance-for-symbolic-regression/108800 "2024-01-14T21:29:48Z")\
**Posts on this page:** 1\
**Showing post:** 10

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [January 15, 2024, 7:56am UTC](https://discourse.julialang.org/t/improving-gpu-performance-for-symbolic-regression/108800/10 "2024-01-15T07:56:07Z")

</div>

Alright so that big red bar in that flame graph turned out to be this call in CUDA.jl in `execution.jl`:

```julia
function (kernel::HostKernel)(args...; threads::CuDim=1, blocks::CuDim=1, kwargs...)
    call(kernel, map(cudaconvert, args)...; threads, blocks, kwargs...)
end

```

If I simply apply this change to it:

```diff
- function (kernel::HostKernel)(args...; threads::CuDim=1, blocks::CuDim=1, kwargs...)
+ function (kernel::HostKernel)(args::Vararg{Any,M}; threads::CuDim=1, blocks::CuDim=1, kwargs...) where {M}

```

this makes it specialize and gets another 20% improvement!

This is because Vararg doesn’t automatically specialize [Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/#Be-aware-of-when-Julia-avoids-specializing) – so that `map(cudaconvert, args)...` call causes a bit of damage.

@maleadt are you okay if we upstream this change?

---

_[View the full topic](https://discourse.julialang.org/t/improving-gpu-performance-for-symbolic-regression/108800)._
