# Matrix handling in memory

**URL:** <https://discourse.julialang.org/t/matrix-handling-in-memory/114258>\
**Category:** Performance\
**Created:** [May 14, 2024, 5:00pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258 "2024-05-14T17:00:45Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 14, 2024, 5:00pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/1 "2024-05-14T17:00:45Z")

</div>

I have a code that is divided in three parts: pre-processing, iterative scheme, and post-processing. In the pre-processing, it calculates a matrix (that can be very large); but it won’t be used in the iterative scheme, which is the most time consuming part of the whole code. This matrix will be used in the post-processing phase only. So, I would like to know from you if there is a way to save this matrix in disc and save memory cost, which one would be the best option? Is there a Julia package ready-to-use for this task?

---

<div class="post-metadata">

**Author:** ![KnutAM](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/knutam/32/37720_2.png) [@KnutAM](https://discourse.julialang.org/u/KnutAM)\
**Post date:** [May 14, 2024, 5:06pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/2 "2024-05-14T17:06:42Z")

</div>

You could use the [`Serialization`](https://docs.julialang.org/en/v1.10/stdlib/Serialization/) standard library’s `serialize` and `deserialize`

E.g.

```julia
using Serialization
M = rand(200,200);
serialize("matrix.bin", M)

x = open("matrix.bin", "r") do io
    deserialize(io)
end;

```

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 14, 2024, 5:15pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/3 "2024-05-14T17:15:48Z")

</div>

Thank you for your answer! Do you have any experience with the performance of Serialization? Have you ever tried the package [GitHub - PetrKryslUCSD/DataDrop.jl: Numbers and matrices and strings stored to disk and retrieved again.](https://github.com/PetrKryslUCSD/DataDrop.jl)?

---

<div class="post-metadata">

**Author:** ![KnutAM](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/knutam/32/37720_2.png) [@KnutAM](https://discourse.julialang.org/u/KnutAM)\
**Post date:** [May 14, 2024, 5:32pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/4 "2024-05-14T17:32:24Z")

</div>

I haven’t benchmarked, but I would guess Serialization is faster or at least comparable

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [May 14, 2024, 6:05pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/5 "2024-05-14T18:05:29Z")

</div>

If the matrix just contain numbers I would use something more portable than Serialization, for example the Arrow format: [User Manual · Arrow.jl](https://arrow.apache.org/julia/stable/manual/#Arrow.write)

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 14, 2024, 6:13pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/6 "2024-05-14T18:13:23Z")

</div>

Thank you! Yes, it is real square matrix.

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 14, 2024, 10:26pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/7 "2024-05-14T22:26:43Z")

</div>

I’ve been testing with Arrow. Do I need to convert the matrix to a DataFrame?

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [May 14, 2024, 10:42pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/8 "2024-05-14T22:42:32Z")

</div>

Here is some simple HDF5.jl usage:

```julia
julia> using HDF5

julia> A = rand(16, 16)
16×16 Matrix{Float64}:
 0.221447 0.811268 0.353941 … 0.9793 0.242755 0.854743
 0.723401 0.720135 0.0715644 0.828959 0.330749 0.343337
 0.699929 0.966771 0.287677 0.325095 0.469268 0.3752
 0.681071 0.927654 0.890345 0.0767709 0.655433 0.447086
 0.754597 0.123622 0.117297 0.776716 0.532075 0.840728
 0.547704 0.297954 0.505034 … 0.931126 0.384665 0.896195
 0.435483 0.434247 0.0798438 0.613545 0.891797 0.890092
 0.817429 0.952932 0.620872 0.808312 0.94135 0.740395
 0.481258 0.6092 0.668251 0.988837 0.438144 0.958993
 0.828436 0.333143 0.438589 0.237257 0.100838 0.576843
 0.652142 0.774052 0.648885 … 0.016766 0.719377 0.215559
 0.364151 0.579178 0.49379 0.885932 0.239334 0.174138
 0.619411 0.850066 0.828862 0.793094 0.534864 0.50797
 0.300589 0.354683 0.224314 0.229821 0.347456 0.397955
 0.0131737 0.616927 0.181855 0.147175 0.615718 0.261567
 0.104203 0.233343 0.89632 … 0.983387 0.0355454 0.62741

julia> h5write("mydata.h5", "A", A)

julia> h5read("mydata.h5", "A")
16×16 Matrix{Float64}:
 0.221447 0.811268 0.353941 … 0.9793 0.242755 0.854743
 0.723401 0.720135 0.0715644 0.828959 0.330749 0.343337
 0.699929 0.966771 0.287677 0.325095 0.469268 0.3752
 0.681071 0.927654 0.890345 0.0767709 0.655433 0.447086
 0.754597 0.123622 0.117297 0.776716 0.532075 0.840728
 0.547704 0.297954 0.505034 … 0.931126 0.384665 0.896195
 0.435483 0.434247 0.0798438 0.613545 0.891797 0.890092
 0.817429 0.952932 0.620872 0.808312 0.94135 0.740395
 0.481258 0.6092 0.668251 0.988837 0.438144 0.958993
 0.828436 0.333143 0.438589 0.237257 0.100838 0.576843
 0.652142 0.774052 0.648885 … 0.016766 0.719377 0.215559
 0.364151 0.579178 0.49379 0.885932 0.239334 0.174138
 0.619411 0.850066 0.828862 0.793094 0.534864 0.50797
 0.300589 0.354683 0.224314 0.229821 0.347456 0.397955
 0.0131737 0.616927 0.181855 0.147175 0.615718 0.261567
 0.104203 0.233343 0.89632 … 0.983387 0.0355454 0.62741

```

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 15, 2024, 12:15am UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/9 "2024-05-15T00:15:50Z")

</div>

I did some benchmarking and found the following results. A random matrix was saved and retrieved by each one of the methods for a number of realizations (sort of Monte Carlo Simulation to obtain the mean time and its standard deviation). Then, I used t-Student cumulative distribution (alpha=0.05) to calculate the confidence interval for each mean and plotted it. As far as the experiments go, serialization seams to be the fastest approach.  
 ![means_methods](https://global.discourse-cdn.com/julialang/original/3X/1/d/1d61775d725b1d04e679b18b5edf7e163611d1cf.png)

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 15, 2024, 12:18am UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/10 "2024-05-15T00:18:14Z")

</div>

[matrix\_handler.jl](https://discourse.julialang.org/uploads/short-url/qB9VQnisTr6coS4jfnOd7qd5XP8.jl) (6.7 KB)

Here is the code I used to do the experiments.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [May 15, 2024, 1:01am UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/11 "2024-05-15T01:01:44Z")

</div>

Oh dear, is this a performance question? I was going for simple and standardized.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [May 15, 2024, 3:48am UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/12 "2024-05-15T03:48:26Z")

</div>

If you don’t use the matrix during the iterative scheme, why not calculate the matrix after the iterative scheme?

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 15, 2024, 11:35am UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/13 "2024-05-15T11:35:45Z")

</div>

This matrix is the product of a Schur decomposition that must be calculated in the pre-processing step.

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [May 15, 2024, 1:10pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/14 "2024-05-15T13:10:12Z")

</div>

Would memory-mapping the matrix be useful here? I am imagining that if the matrix is memory-mapped when it is created but not used during the second stage of the calculation its in-memory image would be used for other purposes then restored during the third stage.

---

<div class="post-metadata">

**Author:** ![MatheusJanczkowski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matheusjanczkowski/32/206380_2.png) [@MatheusJanczkowski](https://discourse.julialang.org/u/MatheusJanczkowski)\
**Post date:** [May 15, 2024, 3:53pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/15 "2024-05-15T15:53:43Z")

</div>

Interesting! Do you have any working example in Julia?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [May 15, 2024, 3:57pm UTC](https://discourse.julialang.org/t/matrix-handling-in-memory/114258/16 "2024-05-15T15:57:15Z")

</div>

```julia
julia> using Mmap

julia> x = Mmap.mmap(Matrix{Float64}, (10_000, 10_000));

```

then just write into it.
