# Efficient Initialization of huge sparse arrays

**URL:** <https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691>\
**Category:** New to Julia\
**Created:** [September 12, 2019, 2:28pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691 "2019-09-12T14:28:22Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [September 12, 2019, 2:28pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/1 "2019-09-12T14:28:22Z")

</div>

There are two ways one can initialize a NXN sparse matrix, whose entries are to be read from one/multiple text files. Which one is faster ? I need the more efficient one, as N is large, typically 10^6.

1. I could store the (x,y) indices in arrays `x`, `y`, the entries in an array `v` and declare  
`K = sparse(x,y,value);`
2. I could declare  
`K = spzeros(N)`  
then read of the (i,j) coordinates and values v and insert them as  
`K[i,j]=v;`  
as they are being read.

I found no tips about this in Julia’s page on sparse arrays.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [September 12, 2019, 2:32pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/2 "2019-09-12T14:32:44Z")

</div>

Don’t insert values one by one: that will be tremendously inefficient since the storage in the sparse matrix needs to be reallocated over and over again.

---

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [September 12, 2019, 2:46pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/3 "2019-09-12T14:46:36Z")

</div>

Thanks, that is exactly the information I needed.

---

<div class="post-metadata">

**Author:** ![rdeits](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rdeits/32/286_2.png) [@rdeits](https://discourse.julialang.org/u/rdeits)\
**Post date:** [September 12, 2019, 3:01pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/4 "2019-09-12T15:01:02Z")

</div>

You can also use BenchmarkTools.jl to verify this:

```julia
julia> using SparseArrays

julia> using BenchmarkTools

julia> I = rand(1:1000, 1000); J = rand(1:1000, 1000); X = rand(1000);

julia> function fill_spzeros(I, J, X)
         x = spzeros(1000, 1000)
         @assert axes(I) == axes(J) == axes(X)
         @inbounds for i in eachindex(I)
           x[I[i], J[i]] = X[i]
         end
         x
       end
fill_spzeros (generic function with 1 method)

julia> @btime sparse($I, $J, $X);
  10.713 μs (12 allocations: 55.80 KiB)

julia> @btime fill_spzeros($I, $J, $X);
  96.068 μs (22 allocations: 40.83 KiB)

```

---

<div class="post-metadata">

**Author:** ![pablosanjose](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pablosanjose/32/7006_2.png) [@pablosanjose](https://discourse.julialang.org/u/pablosanjose)\
**Post date:** [September 12, 2019, 6:28pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/5 "2019-09-12T18:28:09Z")

</div>

Just for completeness, if you need to build/update sparse matrices repeatedly, have a look also at the parent method `SparseArray.sparse!` (not exported)

EDIT: incidentally, you can even avoid allocating, building and processing your `x`, `y`, `value` arrays if you can generate your matrix nonzeros ordered by column. Then, you can build a `SparseMatrixCSC` directly by generating its internal `colptrs`, `rowvals` and `nzval` fields efficiently. Not sure if you want to mess with such details, though I think you can gain quite a bit of efficiency this way

---

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [September 15, 2019, 9:09pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/6 "2019-09-15T21:09:58Z")

</div>

Thank you for the painstaking effort !

---

<div class="post-metadata">

**Author:** ![iamsuddhasattwa](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iamsuddhasattwa/32/7441_2.png) [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Post date:** [September 15, 2019, 9:19pm UTC](https://discourse.julialang.org/t/efficient-initialization-of-huge-sparse-arrays/28691/7 "2019-09-15T21:19:53Z")

</div>

Thanks for the tip !
