# Broadcasting in CUDA kernels

**URL:** <https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521>\
**Category:** General Usage\
**Tags:** gpu, cuda\
**Created:** [June 7, 2021, 2:06pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521 "2021-06-07T14:06:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![fedoroff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fedoroff/32/53209_2.png) [@fedoroff](https://discourse.julialang.org/u/fedoroff)\
**Post date:** [June 7, 2021, 2:06pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/1 "2021-06-07T14:06:41Z")

</div>

Hello,

In the following example code there are two kernels, `kernel1` and `kernel2`, which do exactly the same, but `kernel1` uses broadcasting and, as a result, causes an error. Is it an expected behaviour? Is it possible to use broadcasting in CUDA kernels?

Thank you.

```julia
using CUDA

function kernel1(a, b)
    @. a = b
    return nothing
end

function kernel2(a, b)
    for i=1:length(a)
        a[i] = b[i]
    end
    return nothing
end

a = CUDA.zeros(10)
b = CUDA.ones(10)

# @cuda threads=256 kernel1(a, b)
# ERROR: LoadError: InvalidIRError: compiling kernel kernel1(CuDeviceVector{Float32, 1}, CuDeviceVector{Float32, 1}) resulted in invalid LLVM IR
# Reason: unsupported dynamic function invocation (call to typename(a::DataType) in Base at essentials.jl:304)

@cuda threads=256 kernel2(a, b)

```

---

<div class="post-metadata">

**Author:** ![ozmaden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ozmaden/32/14438_2.png) [@ozmaden](https://discourse.julialang.org/u/ozmaden)\
**Post date:** [June 7, 2021, 2:48pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/2 "2021-06-07T14:48:00Z")

</div>

Broadcasting and kernels can be seen as opposites. Broadcasting eases your life, so that you do not need to write kernels. Broadcasting therefore can not be used inside a kernel, inside a kernel you are processing scalar values. Broadcasting launches generic kernels that julia sees fitting for the current operation you are trying to perform. This would be a broadcasting solution that replaces `kernel1`:

```julia
using CUDA
a = CUDA.zeros(10)
b = CUDA.ones(10)
a .= b

```

And your `kernel2` would be the “Kernel”-based approach to solve the problem. As you can see there is rarely a reason to write your own kernels in CUDA.jl. From my experience it only comes up when you are not satisfied with the performance of the broadcasting solution.

---

<div class="post-metadata">

**Author:** ![fedoroff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fedoroff/32/53209_2.png) [@fedoroff](https://discourse.julialang.org/u/fedoroff)\
**Post date:** [June 7, 2021, 3:02pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/3 "2021-06-07T15:02:26Z")

</div>

In my actual code I have a two-dimensional array and I want to perform a set of operations along one dimension while parallelize along the other. Something like this:

```julia
using CUDA

function kernel1(a, b)
    id = (blockIdx().x - 1) * blockDim().x + threadIdx().x
    stride = blockDim().x * gridDim().x
    N1, N2 = size(a)
    for i=id:stride:N1
        @. a[i, :] = b
    end
    return nothing
end

function kernel2(a, b)
    id = (blockIdx().x - 1) * blockDim().x + threadIdx().x
    stride = blockDim().x * gridDim().x
    N1, N2 = size(a)
    for i=id:stride:N1
        for j=1:N2
            a[i, j] = b[j]
        end
    end
    return nothing
end

N1 = 100
N2 = 10

a = CUDA.zeros((N1, N2))
b = CUDA.ones(N2)

# @cuda threads=N1 kernel1(a, b)

@cuda threads=N1 kernel2(a, b)

```

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [June 7, 2021, 4:12pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/4 "2021-06-07T16:12:57Z")

</div>

For simple cases like this, you can just use plain ol’ broadcasting:

```julia
julia> a .= b'
100×10 CuArray{Float32, 2}:
 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0
 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0
[...]

```

If your actual kernel function can be written without referring to indices, you can broadcast that, too:

```julia
julia> f(x) = 2asin(x)
f (generic function with 1 method)

julia> @. a = f(b')
100×10 CuArray{Float32, 2}:
 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159
 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159 3.14159
[...]

```

---

<div class="post-metadata">

**Author:** ![fedoroff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fedoroff/32/53209_2.png) [@fedoroff](https://discourse.julialang.org/u/fedoroff)\
**Post date:** [June 7, 2021, 4:24pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/5 "2021-06-07T16:24:36Z")

</div>

Thank you, but, unfortunately, my case is not so simple. ☹

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [June 8, 2021, 10:08am UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/6 "2021-06-08T10:08:13Z")

</div>

Broadcast generally isn’t GPU compatible; it might work with e.g. StaticArrays, but generally you don’t want to do so because that means each thread will be iterating a number of elements (whereas you typically then want to use a parallel reduction and share the work across threads, if possible).

---

<div class="post-metadata">

**Author:** ![fedoroff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fedoroff/32/53209_2.png) [@fedoroff](https://discourse.julialang.org/u/fedoroff)\
**Post date:** [June 8, 2021, 12:45pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/7 "2021-06-08T12:45:11Z")

</div>

As I tried to show in my second example, the broadcasts can be useful when one has a deal with multidimensional arrays with parallelization along a specific axis. Another case is when you apply an external function with internal broadcasts to one of the array’s axis.  
Would it be hard to fix this incompatibility by writing a specific broadcast style for CuDeviceArrays? I thought, in the end, the broadcast is replaced by a corresponding loop.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [June 8, 2021, 12:58pm UTC](https://discourse.julialang.org/t/broadcasting-in-cuda-kernels/62521/8 "2021-06-08T12:58:44Z")

</div>

> [@fedoroff](#):
>
> Would it be hard to fix this incompatibility by writing a specific broadcast style for CuDeviceArrays? I thought, in the end, the broadcast is replaced by a corresponding loop.

That should be possible, yes.
