# Trying to understand CUDA benchmark

**URL:** <https://discourse.julialang.org/t/trying-to-understand-cuda-benchmark/73639>\
**Category:** GPU\
**Created:** [December 26, 2021, 9:00am UTC](https://discourse.julialang.org/t/trying-to-understand-cuda-benchmark/73639 "2021-12-26T09:00:58Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jakub\_Mitura](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakub_mitura/32/19496_2.png) [@Jakub\_Mitura](https://discourse.julialang.org/u/Jakub_Mitura)\
**Post date:** [December 26, 2021, 9:00am UTC](https://discourse.julialang.org/t/trying-to-understand-cuda-benchmark/73639/1 "2021-12-26T09:00:58Z")

</div>

I am benchmarking my kernel that is relatively simple

1. it iterates over 2 arrays of the same size and accumulates in local variables times both arrays agree or not
2. warp reduction of local variables
3. thread block reduction using shared memory

What is hard to understand for me is source of huge variability of kernel speed as seen in the image below - all other non system applications are closed - for a context benchmarking is done on RTX 3080 Windows 10 arrays are of size 826×512×512 Array{Float32, 3} - every time the same arrays are used

code

```julia
BenchmarkTools.DEFAULT_PARAMETERS.samples = 500
BenchmarkTools.DEFAULT_PARAMETERS.seconds =600
BenchmarkTools.DEFAULT_PARAMETERS.gcsample = true
@benchmark CUDA.@sync kernelFunction()

```

![image](https://global.discourse-cdn.com/julialang/original/3X/c/5/c59f28a0390d5ad765edd5cb6b5fd2883c0b1736.png)

Similar function in PyTorch is taking on the same data 24 ms - I am trying to get in this range or better - yet there is some problem with benchamrking as I see - what I am doing wrong?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [December 28, 2021, 8:33am UTC](https://discourse.julialang.org/t/trying-to-understand-cuda-benchmark/73639/2 "2021-12-28T08:33:30Z")

</div>

Hard to tell without an MWE. Try running under NSight Systems with `CUDA.@profile`. You can additionally use `NVTX.@range` to mark specific parts of your application and visualize them in the timeline. If you then run multiple invocations it’ll hopefully be clear where the variability comes from.
