# Computing QR decomposition many times in parallel

**URL:** <https://discourse.julialang.org/t/computing-qr-decomposition-many-times-in-parallel/75241>\
**Category:** GPU\
**Tags:** qr\
**Created:** [January 26, 2022, 5:11pm UTC](https://discourse.julialang.org/t/computing-qr-decomposition-many-times-in-parallel/75241 "2022-01-26T17:11:27Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![coviktor](https://avatars.discourse-cdn.com/v4/letter/c/e5b9ba/32.png) [@coviktor](https://discourse.julialang.org/u/coviktor)\
**Post date:** [January 26, 2022, 5:11pm UTC](https://discourse.julialang.org/t/computing-qr-decomposition-many-times-in-parallel/75241/1 "2022-01-26T17:11:27Z")

</div>

I’m new to GPU programming and don’t understand many details. Let’s say I have 2 matrices

```julia
A1 = CUDA.rand(1000,1000)
A2 = CUDA.rand(1000,1000)

```

and I would like to compute their QR decompositions on GPU in parallel (using CUDA.qr). For example, the code below

```julia
Q1,R1 = CUDA.qr(A1)
Q2,R2 = CUDA.qr(A2)

```

does it sequentially. Is there an easy way to do this in parallel?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [January 26, 2022, 6:05pm UTC](https://discourse.julialang.org/t/computing-qr-decomposition-many-times-in-parallel/75241/2 "2022-01-26T18:05:25Z")

</div>

Unlike matrix multiplication, matrix factorizations [aren’t easily parallelizable](https://mrichards.ece.gatech.edu/wp-content/uploads/sites/462/2016/08/Kerr_Campbell_Richards_QRD_on_GPUs.pdf) using GPU hardware, and `CUDA.qr` already exploits device parallelism. For a large number of small input matrices, you may see some benefit by moving to a [batched factorization](https://docs.nvidia.com/cuda/cublas/index.html#cublas-lt-t-gt-geqrfbatched) (available in CUDA.jl as `CUBLAS.geqrf_batched`), but that won’t yield any gains when you only have two large input arrays.

---

<div class="post-metadata">

**Author:** ![coviktor](https://avatars.discourse-cdn.com/v4/letter/c/e5b9ba/32.png) [@coviktor](https://discourse.julialang.org/u/coviktor)\
**Post date:** [January 27, 2022, 4:53pm UTC](https://discourse.julialang.org/t/computing-qr-decomposition-many-times-in-parallel/75241/3 "2022-01-27T16:53:49Z")

</div>

Thanks, batched factorization for 100 of matrices of size 1000x1000 is faster than doing it sequentially using for loop.
