# Large allocations for imported text files

**URL:** <https://discourse.julialang.org/t/large-allocations-for-imported-text-files/78239>\
**Category:** Performance\
**Created:** [March 21, 2022, 7:40pm UTC](https://discourse.julialang.org/t/large-allocations-for-imported-text-files/78239 "2022-03-21T19:40:52Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jdbeck66](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jdbeck66/32/19399_2.png) [@Jdbeck66](https://discourse.julialang.org/u/Jdbeck66)\
**Post date:** [March 21, 2022, 7:40pm UTC](https://discourse.julialang.org/t/large-allocations-for-imported-text-files/78239/1 "2022-03-21T19:40:52Z")

</div>

I’m importing 6 text files that are each between 5 and 8 GB each using BioJulia’s FASTX.

I’m using the code below to read in the data I want to an array for further analysis and processing.

```julia
function getsequences(file,WT)
    open(FASTQ.Reader, file) do reader
        s = []
        q = ExactSearchQuery(dna"N")
        l = length(WT)
        for r in reader
            # Check if no undetermined and long enough
            if FASTX.FASTQ.seqlen(r) ≥ l && !occursin(q, sequence(r))
                push!(s, convert(LongRNA{2},sequence(r))) 
            end 
        end
        return s
    end
end

```

I’m not overly concerned about time - but in looking at the allocations (see below) I’m allocating way more space than the size of the actual files.

 ![Screen Shot 2022-03-21 at 1.30.17 PM](https://global.discourse-cdn.com/julialang/original/3X/c/4/c473fc31b2a55b8ed0c88341429204385316b64e.png)

I’m still trying to get some ideas squared in my mind - Is this because I’m not preallocating space? And how can I do this algorithmically recognizing that file lines and size change - I know there is a way to pass a hint about size that might improve performance? I’m also thinking that this issue comes along with IOing files generally?

Anybody who can share some wisdom on this topic?

JB

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [March 22, 2022, 10:05am UTC](https://discourse.julialang.org/t/large-allocations-for-imported-text-files/78239/2 "2022-03-22T10:05:56Z")

</div>

I know of at least two options to check where allocations originate. First the easy one, use a profiler (I prefer [this one](https://docs.junolab.org/stable/man/juno_frontend/#Profiler)). If you need to know more details you could check [track–allocations](https://docs.julialang.org/en/v1/manual/profile/#Memory-allocation-analysis).

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [March 22, 2022, 12:17pm UTC](https://discourse.julialang.org/t/large-allocations-for-imported-text-files/78239/3 "2022-03-22T12:17:46Z")

</div>

> [@Jdbeck66](#):
>
> ` s = []`

Changing this to an array with a concrete type of element may help. As it seems, you could initialize it with:

```julia
s = LongRNA{2}[]

```

(seems from the code that the type of element that will be added has that type signature).
