# Cuda (Julia vs C++)

**URL:** <https://discourse.julialang.org/t/cuda-julia-vs-c/109865>\
**Category:** Performance\
**Tags:** cuda, benchmark, cudajl\
**Created:** [February 7, 2024, 12:20pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865 "2024-02-07T12:20:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Enlil50](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/enlil50/32/206539_2.png) [@Enlil50](https://discourse.julialang.org/u/Enlil50)\
**Post date:** [February 7, 2024, 12:20pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865/1 "2024-02-07T12:20:09Z")

</div>

Are there any benchmarks between C++ cuda and Cuda.jl?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [February 7, 2024, 12:27pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865/2 "2024-02-07T12:27:47Z")

</div>

We did a comparison against CUDA C with the [Rodinia benchmark suite](https://github.com/JuliaParallel/rodinia) when originally developing CUDA.jl, and the results were good: kernels written in Julia, in the same style as how you would write kernels in C, performs on average pretty much the same. The paper can be found [here](https://arxiv.org/abs/1712.03112), but it’s a couple years old.

There are a couple of things to watch out for: in Julia, array accesses (and many other operations) can throw, introducing additional blocks in the generated code. These can be avoided by e.g. using `@inbounds` being careful about value conversions. Julia also defaults to 64-bit integers, which can make literals and address calculations consume more registers. This generally only matters when micro-optimizing a kernel though.

---

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [February 7, 2024, 12:48pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865/3 "2024-02-07T12:48:56Z")

</div>

> [@maleadt](#):
>
> Julia also defaults to 64-bit integers, which can make literals and address calculations consume more registers.

Is something like this problematic? Because the 1 is a `Int64`?

```julia
out[yrot_int + 1, xrot_int + 1, c] += xdiff_1minus * ydiff_1minus * o

```

---

<div class="post-metadata">

**Author:** ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)\
**Post date:** [February 7, 2024, 3:45pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865/4 "2024-02-07T15:45:23Z")

</div>

> [@roflmaostc](#):
>
> `out[yrot_int + 1, xrot_int + 1, c] += xdiff_1minus * ydiff_1minus * o`

I can’t speak for whether this is a problem in CUDA, but I can suggest `yrot_int + oneunit(yrot_int)`, `yrot_int + oftype(yrot_int, 1)` or `yrot_int + true` as operations that will increment `yrot_int` while preserving its type.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [February 12, 2024, 12:03pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865/5 "2024-02-12T12:03:44Z")

</div>

It would be, however, much of Julia’s indexing infrastructure currently assumes `Int` so your indices would likely get promoted anyway. In [WIP: Add an index typevar to CuDeviceArray. by maleadt · Pull Request #1895 · JuliaGPU/CUDA.jl · GitHub](https://github.com/JuliaGPU/CUDA.jl/pull/1895), I’m experimenting with trying to preserve Int32 indices, but it’s tricky.
