# Batched Matrix solve in CUDA.jl

**URL:** <https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855>\
**Category:** GPU\
**Tags:** blas\
**Created:** [January 4, 2021, 10:49pm UTC](https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855 "2021-01-04T22:49:14Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jpdoane](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jpdoane](https://discourse.julialang.org/u/jpdoane)\
**Post date:** [January 4, 2021, 10:49pm UTC](https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855/1 "2021-01-04T22:49:14Z")

</div>

Is there a way to perform batched matrix inversion/ldiv in CUDA.jl, ie:

```julia
Yi = Ri \ Xi

```

where Ri is NxN, and Yi and Xi are NxM, with N and M small (~12). I need to solve a large number of these systems, e.g. 1 \< i \< 50k. Obviously making multiple GPU calls in a loop will be highly inefficient for small N,M and CUDA provides several relevant batched methods, e.g. cusolverDnpotrsBatched() or cublasgetrsBatched().

I don’t believe the current CUDA.jl implementations of inv or ldiv support these batched calls. Is there a workaround or a roadmap for implementing this in CUDA.jl?

---

<div class="post-metadata">

**Author:** ![Kai\_Xu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kai_xu/32/10187_2.png) [@Kai\_Xu](https://discourse.julialang.org/u/Kai_Xu)\
**Post date:** [January 19, 2021, 12:52am UTC](https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855/2 "2021-01-19T00:52:59Z")

</div>

Is [https://github.com/JuliaGPU/CUDA.jl/blob/master/test/cublas.jl#L1646](https://github.com/JuliaGPU/CUDA.jl/blob/master/test/cublas.jl#L1646) what you are looking for?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [January 19, 2021, 1:48am UTC](https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855/3 "2021-01-19T01:48:55Z")

</div>

Take a look at this discussion: [Accelerate solving many matrix problems - #9 by clinton](https://discourse.julialang.org/t/accelerate-solving-many-matrix-problems/40500/9)

---

<div class="post-metadata">

**Author:** ![RobertGregg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertgregg/32/22105_2.png) [@RobertGregg](https://discourse.julialang.org/u/RobertGregg)\
**Post date:** [February 1, 2023, 7:00pm UTC](https://discourse.julialang.org/t/batched-matrix-solve-in-cuda-jl/52855/4 "2023-02-01T19:00:38Z")

</div>

Sorry to bring up old posts but I can across this when trying to solve a similar problem.

I want to make a quick update on the solution because the line:

```julia
 cuuplo = CUDA.CUBLAS.cublasfill('U')

```

in the `hyperreg()` gave a “not defined” error for me (`UndefVarError`). I replaced it with:

```julia
 cuuplo = CUDA.CUBLAS.CUBLAS_FILL_MODE_UPPER

```

to get it working again.

I also wanted to ask if this is still the go-to solution for batch Matrix solves on the GPU?
