# What is the fastest way to parse a string of numbers into a tuple / struct with different field types?

**URL:** <https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905>\
**Category:** General Usage\
**Created:** [November 28, 2020, 6:23pm UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905 "2020-11-28T18:23:11Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [November 28, 2020, 6:23pm UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/1 "2020-11-28T18:23:11Z")

</div>

Suppose I have a large table whose rows look like

```julia
3 1 2.5 4.0 10.1

```

and I know for sure that the first two columns are `Int`s and the next three are `Float64`s.  
Is there a way to use that information to write a fast parsing function?

So far, I’m coming to something like this:

```julia
function parseline(types, line)
    buf = IOBuffer(line)
    tokens = ntuple(length(types)) do _
        readuntil(buf, ' ')
    end
    map(parse, types, tokens)
end

julia> parseline((Int, Int, Float64, Float64, Float64), "3 1 2.5 4.0 10.1")
(3, 1, 2.5, 4.0, 10.1)

```

which works faster than

```julia
map((t, token)->parse(t, token), types, split(line))

```

Is there a fairly easy way to make the operation faster and type-stable?  
In general, I don’t know how many columns the table has and which types they have until I read the header, so hardcoding the types is not an option.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [November 28, 2020, 6:36pm UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/2 "2020-11-28T18:36:30Z")

</div>

I think this has been already solved by [GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other delimited files in the Julia programming language](https://github.com/JuliaData/CSV.jl)

---

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [November 28, 2020, 7:57pm UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/3 "2020-11-28T19:57:20Z")

</div>

Thanks, that’s almost exactly what I need!

---

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [November 29, 2020, 7:42am UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/4 "2020-11-29T07:42:43Z")

</div>

If I have multiple tables per file (a sequence of frames from an MD simulation), is it possible to make CSV.jl stop reading after a certain number of lines? I’ve tried `CSV.File` with the `limit` keyword but it only seems to affect how many lines will be parsed, not how many will be read.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [November 29, 2020, 8:23am UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/5 "2020-11-29T08:23:44Z")

</div>

> [@Vasily\_Pisarev](#):
>
> make CSV.jl stop reading after a certain number of lines?

Since it implements the Tables.jl interface, you can get an iterator with [`Tables.rows`](https://tables.juliadata.org/stable/#Tables.rows-usage-1) and just collect that as many lines as you need.

---

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [November 29, 2020, 1:51pm UTC](https://discourse.julialang.org/t/what-is-the-fastest-way-to-parse-a-string-of-numbers-into-a-tuple-struct-with-different-field-types/50905/6 "2020-11-29T13:51:52Z")

</div>

Thanks, tested that approach. It seems even slower than readline-split-parse (I guess because `CSV.Row` is not concretely-typed).  
Another problem is that eager parsing of the whole file is not desireable, since a file may contain a long trajectory, and loading it whole may lead to OOM.
