# Distributed/out-of-memory/GPU calculations on sparse matrices

**URL:** <https://discourse.julialang.org/t/distributed-out-of-memory-gpu-calculations-on-sparse-matrices/2792>\
**Category:** General Usage\
**Tags:** gpu\
**Created:** [March 21, 2017, 1:23pm UTC](https://discourse.julialang.org/t/distributed-out-of-memory-gpu-calculations-on-sparse-matrices/2792 "2017-03-21T13:23:32Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Axel\_Gagge](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/axel_gagge/32/859_2.png) [@Axel\_Gagge](https://discourse.julialang.org/u/Axel_Gagge)\
**Post date:** [March 21, 2017, 1:23pm UTC](https://discourse.julialang.org/t/distributed-out-of-memory-gpu-calculations-on-sparse-matrices/2792/1 "2017-03-21T13:23:32Z")

</div>

Hello, I am a junior PhD student in theoretical physics. I am working on a DMRG ([Density matrix renormalization group - Wikipedia](https://en.wikipedia.org/wiki/Density_matrix_renormalization_group)) algorithm in Julia.

A key part of the algorithm is singular value decomposition (SVD) and diagonalization (Lanczos or Arnoldi) on large sparse matrices (MxM where M = d^2\*chi^2, chi=10^2-10^4 and d = 2-10 so that M is of the order 10^6-10^10 in worst case).

It seems to me like I need to make sure that the code can work with arrays which don’t fit in memory, and Blocks.jl or ScaLAPACK bindings looks promising. Ideally, I think it would benefit from parallelization and GPU since at least some operations are embarrasingly parallel.

But I am very new to Julia and have only had a brief introduction to MPI: does anyone have recommendations for which tools and design could be useful for this problem?

Regards,  
Axel Gagge

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [March 22, 2017, 2:08am UTC](https://discourse.julialang.org/t/distributed-out-of-memory-gpu-calculations-on-sparse-matrices/2792/2 "2017-03-22T02:08:07Z")

</div>

I would recommend to take a look at [DistributedArrays.jl](https://github.com/JuliaParallel/DistributedArrays.jl) as a first step and there is also @shashi’s [Dagger.jl](https://github.com/JuliaParallel/Dagger.jl) for out-of-core computations.

cc: @andreasnoack

---

<div class="post-metadata">

**Author:** ![andreasnoack](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andreasnoack/32/27_2.png) [@andreasnoack](https://discourse.julialang.org/u/andreasnoack)\
**Post date:** [March 22, 2017, 2:58am UTC](https://discourse.julialang.org/t/distributed-out-of-memory-gpu-calculations-on-sparse-matrices/2792/3 "2017-03-22T02:58:50Z")

</div>

How many non-zeros do you in the array? If you have a big machine, it is likely that you store the sparse matrix in memory. You might want to benefit from more processors and the vectors generated from Lanczos iterations are also dense. You could take a look at the examples in [GitHub - JuliaParallel/Elemental.jl: Julia interface to the Elemental linear algebra library.](https://github.com/JuliaParallel/Elemental.jl) to see how Elemental’s distributed matrix structures can be used together with an iterative SVD written in Julia. Unfortunately, I don’t have the eigenvalue version ready yet.
