# QR decomposition on large (\>1TB) matrix

**URL:** <https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552>\
**Category:** Julia at Scale\
**Tags:** question\
**Created:** [November 21, 2020, 7:06pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552 "2020-11-21T19:06:42Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![casv2](https://avatars.discourse-cdn.com/v4/letter/c/f6c823/32.png) [@casv2](https://discourse.julialang.org/u/casv2)\
**Post date:** [November 21, 2020, 7:06pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/1 "2020-11-21T19:06:42Z")

</div>

I’d like to perform a QR decomposition on a large (\> 1TB) matrix. Much larger than the available memory on a single node on our HPC.

My idea is to set up a DistributedArray ([GitHub - JuliaParallel/DistributedArrays.jl: Distributed Arrays in Julia](https://github.com/JuliaParallel/DistributedArrays.jl)) or MPIArray ([GitHub - barche/MPIArrays.jl: Distributed arrays based on MPI onesided communication](https://github.com/barche/MPIArrays.jl)) to share the data across a set of nodes to be able to fit the matrix into memory. I’d then like to perform (a distributed) QR decomposition or Ridge Regression. However, it seems these packages do not support a QR decomposition?

I’ve tried using Elemental.jl ([GitHub - JuliaParallel/Elemental.jl: Julia interface to the Elemental linear algebra library.](https://github.com/JuliaParallel/Elemental.jl)) but that seems to spawn processes with copies of the matrix, resulting in OutOfMemory() errors.

Can anyone point me in the right direction to try and solve this problem? Many thank in advance for helping me out!

---

<div class="post-metadata">

**Author:** ![rveltz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rveltz/32/2707_2.png) [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Post date:** [November 21, 2020, 7:25pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/2 "2020-11-21T19:25:27Z")

</div>

> [@casv2](#):
>
> My idea is to set up a DistributedArray ([h](https://github.com/JuliaParallel/DistributedArrays.jl)

Can’t you use a Matrix Free version?

---

<div class="post-metadata">

**Author:** ![boywithacoin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/boywithacoin/32/205435_2.png) [@boywithacoin](https://discourse.julialang.org/u/boywithacoin)\
**Post date:** [December 14, 2023, 4:49pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/3 "2023-12-14T16:49:31Z")

</div>

wow, I have never seen or worked with large matrices. How many rows and cols does this matrix have?

---

<div class="post-metadata">

**Author:** ![amontoison](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amontoison/32/218741_2.png) [@amontoison](https://discourse.julialang.org/u/amontoison)\
**Post date:** [December 14, 2023, 11:07pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/4 "2023-12-14T23:07:00Z")

</div>

I think that `QRMumps` can handle this problem with the Julia interface [QRMumps.jl](https://github.com/JuliaSmoothOptimizers/QRMumps.jl). But you will need to compile a local version with StarPU to enable MPI support. The version precompiled by Yggdrasil doesn’t use StarPU.  
QRMumps, like MUMPS, are tailored for very large problems.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 14, 2023, 11:15pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/5 "2023-12-14T23:15:53Z")

</div>

> [@amontoison](#):
>
> QRMumps, like MUMPS, are tailored for very large problems.

MUMPS is for very large _sparse_ problems.

If you have a large _dense_ matrix, you want a parallel dense-direct library, i.e. something like [Elemental.jl](https://github.com/JuliaParallel/Elemental.jl).

> [@casv2](#):
>
> I’ve tried using Elemental.jl but that seems to spawn processes with copies of the matrix, resulting in OutOfMemory() errors.

Probably you aren’t using it correctly? Handling distributed matrices, with only a chunk of the matrix in each process’s memory, is the whole point of the Elemental library AFAIK.

> [@rveltz](#):
>
> Can’t you use a Matrix Free version?

More generally, the question is where does your matrix come from, and can you exploit some special structure? e.g. is it sparse, or is there a fast way to multiply matrix-times-vector? Do you need the _whole_ QR decomposition, or can you use some approximation? e.g. if you are using it to solve a least-squares problem, can you use [randomized least squares](https://www.sciencedirect.com/science/article/abs/pii/S1877750316301508)?

---

<div class="post-metadata">

**Author:** ![abuttari](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abuttari/32/4165_2.png) [@abuttari](https://discourse.julialang.org/u/abuttari)\
**Post date:** [December 17, 2023, 9:17pm UTC](https://discourse.julialang.org/t/qr-decomposition-on-large-1tb-matrix/50552/6 "2023-12-17T21:17:58Z")

</div>

Hello,  
qr\_mumps is actually developed for solving sparse problems through a multifrontal factorization but also contains some dense linear algebra routines including a parallel QR factorization. The issue, though, is that in order to handle a 1TB matrix you need distributed memory parallelism, which qr\_mumps does not support at the moment. One option is to use ScaLAPACK; there is a related discussion [here](https://discourse.julialang.org/t/how-to-load-scalapack-library-from-intel-mkl-library-inter-dependencies/81239).
