# Why the Floating-Point Calculation Efficiency of CUDA.jl Does Not Reach the Official Theoretical Value

**URL:** <https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474>\
**Category:** GPU\
**Created:** [February 2, 2025, 3:09pm UTC](https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474 "2025-02-02T15:09:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![ajsagit](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ajsagit/32/16017_2.png) [@ajsagit](https://discourse.julialang.org/u/ajsagit)\
**Post date:** [February 2, 2025, 3:09pm UTC](https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474/1 "2025-02-02T15:09:29Z")

</div>

```julia
using CUDA
using BenchmarkTools
using Printf

function benchmark_floating_point(T::Type, size=2048)
    aligned_size = div(size, 8) * 8
    A = CUDA.rand(T, (aligned_size, aligned_size))
    B = CUDA.rand(T, (aligned_size, aligned_size))
    C = similar(A)

    CUDA.@sync A * B

    elapsed_time = @elapsed CUDA.@sync C = A * B

    operations = 2 * aligned_size^3
    flops = operations / elapsed_time / 1e12 # TFLOPS
    data_transferred = 3 * sizeof(T) * aligned_size^2 / 1e9 # GB
    bandwidth = data_transferred / elapsed_time # GB/s

    return (elapsed_time, flops, bandwidth)
end

function benchmark_tensor_core(size=2048)
    aligned_size = div(size, 8) * 8

    A = CUDA.rand(Float16, (aligned_size, aligned_size))
    B = CUDA.rand(Float16, (aligned_size, aligned_size))
    C = CUDA.zeros(Float32, (aligned_size, aligned_size))

    CUDA.@sync A * B

    elapsed_time = @elapsed CUDA.@sync C = A * B

    operations = 2 * aligned_size^3
    flops = operations / elapsed_time / 1e12
    return (elapsed_time, flops)
end

function main()
    sizes = [256, 512, 1024, 2048]
    precisions = [Float16, Float32, Float64]

    println("RTX 3080 Ti FP performance (Julia + CUDA.jl)")
    println("=============================================")

    # 测试原生精度
    for T in precisions
        println("\n[FP] ", T)
        for size in sizes
            time, flops, bw = benchmark_floating_point(T, size)
            @printf("Size: %4d x %4d | Time: %.4f s | TFLOPS: %6.2f | Bandwidth: %6.1f GB/s\n",
                    size, size, time, flops, bw)
        end
    end

    println("\n[Tensor Core]")
    ENV["JULIA_CUDA_USE_TENSOR_CORES"] = "1"  
    for size in sizes
        time, flops = benchmark_tensor_core(size)
        @printf("Size: %4d x %4d | Time: %.4f s | TFLOPS: %6.2f (Tensor Core)\n",
                size, size, time, flops)
    end
end

main()

```

```julia
RTX 3080 Ti FP (Julia + CUDA.jl)
=============================================

Float16
Size: 256 x 256 | Time: 0.0000 s | TFLOPS: 1.28 | Bandwidth: 15.0 GB/s
Size: 512 x 512 | Time: 0.0000 s | TFLOPS: 10.09 | Bandwidth: 59.1 GB/s
Size: 1024 x 1024 | Time: 0.0000 s | TFLOPS: 52.76 | Bandwidth: 154.6 GB/s
Size: 2048 x 2048 | Time: 0.0002 s | TFLOPS: 87.74 | Bandwidth: 128.5 GB/s

Float32
Size: 256 x 256 | Time: 0.0000 s | TFLOPS: 1.23 | Bandwidth: 28.9 GB/s
Size: 512 x 512 | Time: 0.0000 s | TFLOPS: 6.74 | Bandwidth: 79.0 GB/s
Size: 1024 x 1024 | Time: 0.0001 s | TFLOPS: 14.82 | Bandwidth: 86.8 GB/s
Size: 2048 x 2048 | Time: 0.0008 s | TFLOPS: 21.84 | Bandwidth: 64.0 GB/s

Float64
Size: 256 x 256 | Time: 0.0001 s | TFLOPS: 0.27 | Bandwidth: 12.5 GB/s
Size: 512 x 512 | Time: 0.0007 s | TFLOPS: 0.37 | Bandwidth: 8.6 GB/s
Size: 1024 x 1024 | Time: 0.0050 s | TFLOPS: 0.43 | Bandwidth: 5.1 GB/s
Size: 2048 x 2048 | Time: 0.0367 s | TFLOPS: 0.47 | Bandwidth: 2.7 GB/s

[Tensor Core]
Size: 256 x 256 | Time: 0.0000 s | TFLOPS: 1.46 (Tensor Core)
Size: 512 x 512 | Time: 0.0000 s | TFLOPS: 8.66 (Tensor Core)
Size: 1024 x 1024 | Time: 0.0001 s | TFLOPS: 21.80 (Tensor Core)
Size: 2048 x 2048 | Time: 0.0002 s | TFLOPS: 74.70 (Tensor Core)

```

 ![230900](https://global.discourse-cdn.com/julialang/original/3X/7/7/77ca06d73a8f149cdeb876bdd9ee1e4d445187c4.jpeg)

---

<div class="post-metadata">

**Author:** ![eldee](https://avatars.discourse-cdn.com/v4/letter/e/b5a626/32.png) [@eldee](https://discourse.julialang.org/u/eldee)\
**Post date:** [February 2, 2025, 3:55pm UTC](https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474/2 "2025-02-02T15:55:10Z")

</div>

I’m no expert, but I’d imagine for the first eight values AIDA64 uses benchmarks which are either completely memory-bound, or completely compute-bound. Matrix multiplication does not seem an appropriate test for either. Additionally note that AIDA64’s results will still not be the same as the (official) theoretical values.

Also, in

> [@ajsagit](#):
>
> ```julia
> C = similar(A)
> elapsed_time = @elapsed CUDA.@sync C = A * B
> 
> ```

the declared `C` matrix is not used. Instead `A * B` creates a new matrix, which we also call `C`. You should use something like `mul!(C, A, B)` from LinearAlgebra.jl.

> [@ajsagit](#):
>
> ```julia
> operations = 2 * aligned_size^3
> 
> ```

I’m not convinced this is fully (or sufficiently) accurate. You’re probably better off testing something more straightforward than (optimised) matrix multiplication.
