# Reading tab-delimited file & memory allocation

**URL:** <https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727>\
**Category:** New to Julia\
**Tags:** memory-allocation, io\
**Created:** [February 18, 2022, 9:51pm UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727 "2022-02-18T21:51:44Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![lln](https://avatars.discourse-cdn.com/v4/letter/l/ecae2f/32.png) [@lln](https://discourse.julialang.org/u/lln)\
**Post date:** [February 18, 2022, 9:51pm UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/1 "2022-02-18T21:51:44Z")

</div>

Hi, I have the following code for parsing a tab-delimited file

```julia
function read_file(file::String, chrom::Int64)::Dict
    vld_dict::Dict{String, Vector{String}} = Dict("SNP" => String[], "A1" => String[], "A2" => String[])
    nsnps::Int64 = 0
                
    open(file) do f
        while ! eof(f)
            ll = readline(f) .|> s -> split(s, "\t")
            if parse(Int64, ll[1]) == chrom
                nsnps += 1
                push!(vld_dict["SNP"], ll[2])
                push!(vld_dict["A1"], ll[5])
                push!(vld_dict["A2"], ll[6])
            end
        end
    end
    return vld_dict
end

```

Calling it using `@time read_file("test.bim", 22)` gives  
`0.001816 seconds (7.52 k allocations: 496.266 KiB)` (for the second run). Is there a way to optimize this to allocate less memory? The file `test.bim` only has 1000 lines that are `chrom = 22`

Thanks!

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 18, 2022, 10:15pm UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/2 "2022-02-18T22:15:22Z")

</div>

> [@lln](#):
>
> `ll = readline(f) .|> s -> split(s, "\t")`

A small optimization is to avoid pipes, as they usually allocate more than explicit writing:

```julia
ll = split(readline(f) , "\t")

```

---

<div class="post-metadata">

**Author:** ![lln](https://avatars.discourse-cdn.com/v4/letter/l/ecae2f/32.png) [@lln](https://discourse.julialang.org/u/lln)\
**Post date:** [February 18, 2022, 10:19pm UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/3 "2022-02-18T22:19:58Z")

</div>

Thanks! That does improve things a bit (`0.001276 seconds (6.52 k allocations: 480.641 KiB)`

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [February 19, 2022, 12:27am UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/4 "2022-02-19T00:27:55Z")

</div>

Are you using a CSV parser like [CSV.jl](https://github.com/JuliaData/CSV.jl)? Even if you don’t use it, that code can help you understand how to write a fast parser.

One challenge is that Julia does make it a bit harder than it could to avoid building lots of new strings instead of reusing a fixed byte buffer. That’s one of the things optimized parsers handle.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [February 19, 2022, 12:34am UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/5 "2022-02-19T00:34:11Z")

</div>

> [@johnmyleswhite](#):
>
> building lots of new strings instead of reusing a fixed byte buffer

Wouldn’t a `split!(buff, string, char)` function be a nice addition to base?

(and a `readline!(buff, f)` as well).

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [February 19, 2022, 11:58am UTC](https://discourse.julialang.org/t/reading-tab-delimited-file-memory-allocation/76727/6 "2022-02-19T11:58:14Z")

</div>

Yes, but I would want something larger scoped: something that lets me use all of the core string functions, but operating on a byte buffer I control and operate on via mutation. And in Base so people don’t forget to add methods to it in the future 🙂

It would be like the never ratified [`string_view`](https://devblogs.microsoft.com/cppblog/stdstring_view-the-duct-tape-of-string-types/) for C++.
