# How to broadcast or batch multiply a "batch" of matrices with another matrix on the GPU?

**URL:** <https://discourse.julialang.org/t/how-to-broadcast-or-batch-multiply-a-batch-of-matrices-with-another-matrix-on-the-gpu/67259>\
**Category:** GPU\
**Created:** [August 29, 2021, 2:54am UTC](https://discourse.julialang.org/t/how-to-broadcast-or-batch-multiply-a-batch-of-matrices-with-another-matrix-on-the-gpu/67259 "2021-08-29T02:54:46Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 29, 2021, 2:54am UTC](https://discourse.julialang.org/t/how-to-broadcast-or-batch-multiply-a-batch-of-matrices-with-another-matrix-on-the-gpu/67259/1 "2021-08-29T02:54:46Z")

</div>

I have a 3 dimension tensor, where each “slice” of the tensor is a matrix. I want to multiple each slice by another matrix and store there result in a 3D tensor/array.

How do I do that in the most efficient way using the GPU?

E.g. I have

```julia
using CUDA
CUDA.allowscalar(false)

tensor = rand(4, 4, 1000) |> cu
matrix = rand(4,4) |> cu

result = mapslices( slice-> slice*matrix, tensor, dims=(1,2))

```

This fails due to scalar indexing not allowed.

Do I need to write a kernel myself?

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [August 29, 2021, 3:56am UTC](https://discourse.julialang.org/t/how-to-broadcast-or-batch-multiply-a-batch-of-matrices-with-another-matrix-on-the-gpu/67259/2 "2021-08-29T03:56:06Z")

</div>

No, you need `⊠`:

```julia
julia> using NNlib, NNlibCUDA

julia> result ≈ batched_mul(tensor, matrix)
true

```

Note that if you wanted the other way around, then you could also do it by reshaping:

```julia
julia> using TensorCore

julia> @btime CUDA.@sync $matrix ⊡ $tensor;
  33.002 μs (34 allocations: 784 bytes)

julia> @btime CUDA.@sync $matrix ⊠ $tensor;
  27.840 μs (12 allocations: 368 bytes)

julia> matrix ⊠ tensor ≈ matrix ⊡ tensor ≈ mapslices(slice-> matrix * slice, tensor, dims=(1,2))
true # on CPU

```
