# Batched Matrix Multiply

**URL:** <https://discourse.julialang.org/t/batched-matrix-multiply/42332>\
**Category:** General Usage\
**Tags:** gpu, blas, linearalgebra, cuarrays\
**Created:** [July 1, 2020, 2:11am UTC](https://discourse.julialang.org/t/batched-matrix-multiply/42332 "2020-07-01T02:11:31Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [July 1, 2020, 7:53am UTC](https://discourse.julialang.org/t/batched-matrix-multiply/42332/5 "2020-07-01T07:53:42Z")

</div>

Here are a couple of alternatives. If you use a vector of matrices, that’s faster, otherwise you can get the same effect with `eachslice`, though at some performance cost. And of course, it’s much faster with StaticArrays.

```julia
using BenchmarkTools, Test

A_ = [rand(4,3) for _ in 1:2];
B_ = [rand(3,4) for _ in 1:2];
A = cat(A_...; dims=3)
B = cat(B_...; dims=3)

foo(X, Y) = X .* Y
bar(X, Y) = eachslice(X; dims=3) .* eachslice(Y; dims=3)

```

```julia
julia> @test foo(A_, B_) == bar(A, B)
Test Passed

julia> @btime foo($A_, $B_)
  540.212 ns (3 allocations: 512 bytes)

julia> @btime bar($A, $B);
  1.797 μs (17 allocations: 1.11 KiB)

```

---

_[View the full topic](https://discourse.julialang.org/t/batched-matrix-multiply/42332)._
