# Best practice for reading large matrix with repeated spaces as text

**URL:** <https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685>\
**Category:** General Usage\
**Tags:** dataframes, csv\
**Created:** [April 20, 2023, 2:43am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685 "2023-04-20T02:43:16Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![sparrowhawk](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sparrowhawk/32/9148_2.png) [@sparrowhawk](https://discourse.julialang.org/u/sparrowhawk)\
**Post date:** [April 20, 2023, 2:43am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/1 "2023-04-20T02:43:16Z")

</div>

Hi, I have a moderately large text file with `1305639×114` entires in a matrix, separated by repeated whitespaces (not always fixed width) per line. About 1.5 GB or so with the spaces.

On julia 1.8.2 The file reads fine with vanilla `readdlm` from `DelimitedFiles.jl` but takes 100 s.

```julia
@time A = readdlm("myfile.dat", Float64)
106.114928 seconds (446.89 M allocations: 14.807 GiB, 1.00% gc time, 0.67% compilation time)
1305639×114 Matrix{Float64}:
.

```

But because this is the first step on a cluster with many CPUs already running idle, this wastes a lot of CPU time. I have found the following using `DataFrames.jl` and `CSV.jl` a whole lot faster

```julia
A = Array{Float64, 2}(CSV.File("myfile.dat"; ignorerepeated=true, types=Float64, header=false, delim=' ')|>DataFrame)
42.190986 seconds (470.89 k allocations: 3.345 GiB, 35.74% compilation time)
1305639×114 Matrix{Float64}
.

```

But I don’t want to add CSV and DataFrames as package dependencies as this means more dependency alertness on my part. Or am I just being finicky / is there a better way than shown above? Thanks.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 20, 2023, 2:56am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/2 "2023-04-20T02:56:59Z")

</div>

One thing worth mentioning is that if you can avoid working with test files for matrices you can expect pretty major speedups. Binary formats such as Arrow are often ~100x faster to read and write.

---

<div class="post-metadata">

**Author:** ![sparrowhawk](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sparrowhawk/32/9148_2.png) [@sparrowhawk](https://discourse.julialang.org/u/sparrowhawk)\
**Post date:** [April 20, 2023, 3:09am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/3 "2023-04-20T03:09:34Z")

</div>

I absolutely agree, but unfortunately I need to write something that reads a horrible legacy format of instrument data for a whole lot of people. I guess one thing to do could be read in text and write out as a binary before I start the cluster job … but would like to avoid duplicating input data if possible.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 20, 2023, 3:33am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/4 "2023-04-20T03:33:17Z")

</div>

if you are doing to read it more than once, I’d recommend reading as text and writing binary first. otherwise I’d just use the csv solution.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 20, 2023, 6:58am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/5 "2023-04-20T06:58:32Z")

</div>

> [@Oscar\_Smith](#):
>
> I’d just use the csv solution.

To add to this. You just need CSV.jl. You do not need to add DataFrames.jl as a dependency.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [April 20, 2023, 8:35am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/6 "2023-04-20T08:35:12Z")

</div>

Here’s how you might do it without DataFrames:

```julia
julia> using CSV, DelimitedFiles

julia> writedlm("myfile.dat", rand(1_300_000, 114));

julia> @time Tables.matrix((CSV.File("myfile.dat"; ignorerepeated = true, types = Float64, header = false, delim = '\t')));
 15.626529 seconds (48.21 k allocations: 2.209 GiB, 0.59% gc time)

```

this isn’t the first call so I’ve taken out the compilation overhead (although I’m on Julia 1.9 so TTFX shouldn’t be a big issue anyway). ((Also the above writes Float64 numbers so the file is 2.6 rather than 1.5 GB, but that won’t make a huge difference to the timings))

---

<div class="post-metadata">

**Author:** ![sparrowhawk](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sparrowhawk/32/9148_2.png) [@sparrowhawk](https://discourse.julialang.org/u/sparrowhawk)\
**Post date:** [April 20, 2023, 8:49am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/7 "2023-04-20T08:49:38Z")

</div>

I see, but could you please clarify? I can’t do

```julia
julia> using CSV

julia> data = 
       """
       1.0 0.0 1.0
       0.0 1.0 0.0
       0.0 0.0 1.0
       """
"1.0 0.0 1.0\n0.0 1.0 0.0\n0.0 0.0 1.0\n"

julia> CSV.File(IOBuffer(data), header=false)|>DataFrame|>Matrix
ERROR: UndefVarError: DataFrame not defined
Stacktrace:
 [1] top-level scope
   @ REPL[3]:1

```

as this does not work if I don’t do a `using DataFrames` first.

---

<div class="post-metadata">

**Author:** ![sparrowhawk](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sparrowhawk/32/9148_2.png) [@sparrowhawk](https://discourse.julialang.org/u/sparrowhawk)\
**Post date:** [April 20, 2023, 8:57am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/8 "2023-04-20T08:57:46Z")

</div>

This is indeed fast. However, I seem to need to do `using Tables` to get this to work on Julia 1.8.2. Won’t that add Tables as a dependency then?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [April 20, 2023, 9:00am UTC](https://discourse.julialang.org/t/best-practice-for-reading-large-matrix-with-repeated-spaces-as-text/97685/9 "2023-04-20T09:00:02Z")

</div>

Ah sorry, I think I had already done `using DataFrames` in the session before I tried this. `Tables` is already a dependency of `CSV` but not re-exported so you’d have to do `CSV.Tables.matrix`.
