# Why is 'trace' slow on CUDA matrices?

**URL:** <https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723>\
**Category:** GPU\
**Tags:** gpu, linearalgebra\
**Created:** [February 16, 2023, 1:51pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723 "2023-02-16T13:51:22Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![HenriDeh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrideh/32/8316_2.png) [@HenriDeh](https://discourse.julialang.org/u/HenriDeh)\
**Post date:** [February 16, 2023, 1:51pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723/1 "2023-02-16T13:51:22Z")

</div>

Hello,

I am kind of surprised to see this benchmark:

```julia
julia> using BenchmarkTools, CUDA, LinearAlgebra

julia> A = rand(128, 128);

julia> Ad = cu(A);

julia> @btime tr($A)
  70.092 ns (0 allocations: 0 bytes)
64.9944482344768

julia> @btime tr($Ad)
  84.400 μs (76 allocations: 3.72 KiB)
64.994446f0

```

Computing the trace of a matrix is 1000 times slower with CUDA. Is there any way around this?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [February 16, 2023, 2:39pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723/2 "2023-02-16T14:39:55Z")

</div>

This is just a memory latency benchmark and CPUs have lower memory latency.

---

<div class="post-metadata">

**Author:** ![HenriDeh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrideh/32/8316_2.png) [@HenriDeh](https://discourse.julialang.org/u/HenriDeh)\
**Post date:** [February 16, 2023, 2:46pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723/3 "2023-02-16T14:46:49Z")

</div>

So if I use traces of CuMatrices in my subroutines it will not create a performance dip?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [February 16, 2023, 3:00pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723/4 "2023-02-16T15:00:37Z")

</div>

> [@HenriDeh](#):
>
> it will not create a performance dip?

it will if your real use case is dominated by things like this example

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [February 16, 2023, 4:05pm UTC](https://discourse.julialang.org/t/why-is-trace-slow-on-cuda-matrices/94723/5 "2023-02-16T16:05:09Z")

</div>

128x128 inputs are tiny; the time it costs to just launch a kernel is about 20us, and (our current implementation of) `tr` requires two kernels. But even with larger inputs the GPU won’t be faster here, as the hardware needs some computational complexity to hide memory latency. You’re essentially doing no compute at all, hence you’re just benchmarking the memory.
