# Fastest way to add arrays

**URL:** https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593
**Category:** Performance
**Tags:** cuda
**Created:** [December 13, 2022, 10:21am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593 "2022-12-13T10:21:51Z")
**Posts on this page:** 13
**Page:** 1

<div class="post-metadata">

### Author: ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)
#### Post date: [December 13, 2022, 10:21am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/1 "2022-12-13T10:21:51Z")

</div>

Hey Julianners,  
Can you help me on how to create fast array addition?  
I believe this is the fastest way:

```julia
using CUDA
using BenchmarkTools
BenchmarkTools.DEFAULT_PARAMETERS.seconds = 1.00
add_cab(c, a, b) = @inbounds begin
  I = (blockIdx().x - 1) * blockDim().x + threadIdx().x
  I > size(c, 1) && return
	ca = CUDA.Const(a)
	cb = CUDA.Const(b)
	c[I] += ca[I] + cb[I] 
	nothing
end
nth=512
N=1000_000               
CG = CUDA.ones(Float32, N)  
AG = CUDA.ones(Float32, N)  
BG = CUDA.ones(Float32, N)  
ITER_GFLOPS = N * 2 / 1000_000_000
@sync @cuda threads=nth blocks=cld(N, nth) add_cab(CG, AG, BG)
b = @belapsed @sync @cuda threads=nth blocks=cld(N, nth) add_cab(CG, AG, BG)
println(ITER_GFLOPS / b) # 3090TI: 387.48 GFlops (Basically this very much sound like a bandwidth problem (the theoretical 40.000Gflops is about 400GFlops).)

```

Instead of 40.000Gflops it is 387Gflops.  
Or do I do something wrong?

[https://www.reddit.com/r/CUDA/comments/zksbpx/fastest\_way\_to\_add\_arrays/](https://www.reddit.com/r/CUDA/comments/zksbpx/fastest_way_to_add_arrays/)

(Note: I know that with FMA I can reach 2x and also with using registry and run the calculation in a for loop 200 times I could reach better speed, but I am not interested in those scenarios.)

---

<div class="post-metadata">

### Author: ![Sixzero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sixzero/32/28737_2.png) [@Sixzero](https://discourse.julialang.org/u/Sixzero)
#### Post date: [December 13, 2022, 10:44am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/2 "2022-12-13T10:44:32Z")

</div>

As a general answer by ChatGPT, sry didn’t want to increase the spam on the discussion: 😃

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/c/7c79cecd19d65925d3d32323a480193ff15c7199.png)

Its usually just bullshit what it can generate.

---

<div class="post-metadata">

### Author: ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)
#### Post date: [December 13, 2022, 10:49am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/3 "2022-12-13T10:49:51Z")

</div>

I don’t think the bottleneck is compute, but rather to the memory bandwidth. The 3090Ti has about 1000 GB/s, IIRC, so divided by 4 bytes per Float32, and 3 read/writes per operation, that gives about 80 Gflops if nothing is in cache, but higher if (parts of) the arrays are already in cache, which depends on surrounding code and size of arrays.

---

<div class="post-metadata">

### Author: ![Sixzero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sixzero/32/28737_2.png) [@Sixzero](https://discourse.julialang.org/u/Sixzero)
#### Post date: [December 13, 2022, 10:54am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/4 "2022-12-13T10:54:11Z")

</div>

I tried some other ideas, with chatGPT, and the two things it added are the use of TensorCores, and compiler optimization to -O3, but idk if the latter matters on GPU.  
Also, it notes that TensorCores are sometimes not possible to be used, in this simple example we could use that, but its pretty tedious to use.

Also as for tensorcores this is the example code I found on the internet, which pretty much doesn’t work. 😃

```julia
using CUDA

# Generate input matrices
a = rand(Float16, (16, 16))
a_dev = CuArray(a)
b = rand(Float16, (16, 16))
b_dev = CuArray(b)
c = rand(Float32, (16, 16))
c_dev = CuArray(c)

# Allocate space for result
d_dev = similar(c_dev)

# Matrix multiply-accumulate kernel (D = A * B + C)
function kernel(a_dev, b_dev, c_dev, d_dev)
    a_frag = WMMA.llvm_wmma_load_a_col_m16n16k16_stride_f16(pointer(a_dev), 16)
    b_frag = WMMA.llvm_wmma_load_b_col_m16n16k16_stride_f16(pointer(b_dev), 16)
    c_frag = WMMA.llvm_wmma_load_c_col_m16n16k16_stride_f32(pointer(c_dev), 16)

    d_frag = WMMA.llvm_wmma_mma_col_col_m16n16k16_f32_f32(a_frag, b_frag, c_frag)

    WMMA.llvm_wmma_store_d_col_m16n16k16_stride_f32(pointer(d_dev), d_frag, 16)
    return
end

@cuda threads=32 kernel(a_dev, b_dev, c_dev, d_dev)

```

---

<div class="post-metadata">

### Author: ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)
#### Post date: [December 13, 2022, 11:36am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/5 "2022-12-13T11:36:44Z")

</div>

@Sixzero I love the answer, but does it gives better speed?

> I don’t think the bottleneck is compute, but rather to the memory bandwidth. The 3090Ti has about 1000 GB/s, IIRC, so divided by 4 bytes per Float32, and 3 read/writes per operation, that gives about 80 Gflops if nothing is in cache, but higher if (parts of) the arrays are already in cache, which depends on surrounding code and size of arrays.

Yes It should be memory bandwidth bottleneck, my problem is that, it is extremly slow compared to the marketed 40.000GFlops theoretical value. And feels totally misleading in the end as basically for a single addition cannot be faster than 1% of the theoretical speed.  
Or am I missing something basically, that is my question…

---

<div class="post-metadata">

### Author: ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)
#### Post date: [December 13, 2022, 11:43am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/6 "2022-12-13T11:43:24Z")

</div>

Just to make sure of the memory boundedness of the calculation, can you replace one or both of the arrays with computed counters (i.e. no memory transfer needed) and see the performance increase?

---

<div class="post-metadata">

### Author: ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)
#### Post date: [December 13, 2022, 12:05pm UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/7 "2022-12-13T12:05:48Z")

</div>

The problem using temporary values is that, even with that, you still have to write into the appropriate memory. So I don’t understand how could I improve the memory usage or something, to minimize the memory bandwidth. I tried to use strides… and many thing… but I couldn’t reach higher GFlops… It is disappointing.

So I had some hope in better cache usage or something that I just cannot achieve. Doesn’t this simple problem can be faster?  
Or can shared\_memory help us in any certain way? I guess it is useful at problems like reduce.

If we could create the fastest addition here we could improve our codes everywhere based on this one good example.

---

<div class="post-metadata">

### Author: ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)
#### Post date: [December 13, 2022, 12:37pm UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/8 "2022-12-13T12:37:13Z")

</div>

gpus have teired memory so the answer is to not optimize a single addition but use a kennel that does more work.

---

<div class="post-metadata">

### Author: ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)
#### Post date: [December 13, 2022, 12:56pm UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/9 "2022-12-13T12:56:39Z")

</div>

The problem is that to add a specified array value, you need to “read” it and the bandwidth is limited.

So the question is how can I use the bandwidth better or something.

---

<div class="post-metadata">

### Author: ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)
#### Post date: [December 13, 2022, 12:59pm UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/10 "2022-12-13T12:59:00Z")

</div>

what are you doing before and after the array add? thinking in terms of kernels rather than vectorized functions is necessary here.

---

<div class="post-metadata">

### Author: ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)
#### Post date: [December 13, 2022, 1:46pm UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/11 "2022-12-13T13:46:03Z")

</div>

I do mean() and many more custom transformation. And also need to use the value of the summations and special structs, so it is barely possible to fuse it (in spite of there is a very bad and ugly way to do it to use only 1 block and do some array flattening).

So that isn’t viable for me, and it is really slow that way.  
I am interested about weather we can make this addition faster so I could make 10-100x speed up if we can use the caching effectively or something.

---

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [December 14, 2022, 7:33am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/12 "2022-12-14T07:33:14Z")

</div>

There’s no magical way to “better use memory”; your kernel is purely reading from and writing to global memory, so is going to be horribly memory bound. The theoretical GFlops are irrelevant here. You’re supposed to increase arithmetic intensity, as @Oscar_Smith mentions, e.g. by fusing kernels together.

---

<div class="post-metadata">

### Author: ![Sixzero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sixzero/32/28737_2.png) [@Sixzero](https://discourse.julialang.org/u/Sixzero)
#### Post date: [December 14, 2022, 11:37am UTC](https://discourse.julialang.org/t/fastest-way-to-add-arrays/91593/13 "2022-12-14T11:37:54Z")

</div>

Yeah, and memory access in a coalescable way seems like the only other thing that could still matter.

I don’t know if shared memory or things like that possible to be used in any way, would be curious if someone could link some good example usage for them! 😇
