# Using readlines with large text files

**URL:** <https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341>\
**Category:** General Usage\
**Created:** [October 15, 2024, 5:20pm UTC](https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341 "2024-10-15T17:20:39Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Colin\_Beckingham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/colin_beckingham/32/4878_2.png) [@Colin\_Beckingham](https://discourse.julialang.org/u/Colin_Beckingham)\
**Post date:** [October 15, 2024, 5:20pm UTC](https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341/1 "2024-10-15T17:20:39Z")

</div>

I’m trying to read a genome string file into Julia using `readlines()` and getting some very strange results. The genome file consists of a gunzipped text file with a descriptive first line followed by a large number of strings of length 60 characters consisting of only ‘A’,‘G’,‘C’ and ‘T’. It should not contain a ‘Y’ and if I read the file into joe editor in Bash and perform a search there are no ‘Y’ characters.

So I read this file into Julia using:

```julia
function opengenome(fafile::String)
	pot = open(fafile)
	potgen = readlines(pot)
	return potgen[2:end]
end

```

The head of the result is like:

```julia
gen = opengenome(genome_athal)
1994471-element Vector{String}:
 "CCCTAAACCCTAAACCCTAAACCCTAAACCTCTGAATCCTTAATCCCTAAATCCCTAAAT"
 "CTTTAAATCCTACATCCATGAATCCCTAAATACCTAATTCCCTAAACCCGAAACCGGTTT"
 "CTCTGGTTGAAAATCATTGTGTATATAATGATAATTTTATCGTTTTTATGTAATTGCTTA"
 "TTGTTGTGTGTAGATTTTTTAAAAATATCATTTGAGGTCAATACAAATCCTATTTCTTGT"

```

and all looks good.

But there is a problem. If I now do a search for ‘Y’ like so:

```julia
for i in 1:length(gen);
 if count('Y',gen[i]) > 0; 
    println(gen[i]); 
 end; 
end

```

The head of the output (there are 86 occurrences in the stored vector) is like:

```julia
GGTGACAAAGTTCCCGGCCAGTGYGTTTGCGGGTAACGACTGTCTTTGTGGCTCTCCACT
NNNNNNNNNNNNNNMTTWTKKCSNYTCYASTTWTTKMRWYTSWAKGWTWWWMWAMWSAWY
AAKWMAMWWWRSAYTAMRWMAAYWYRAACCAMGMWWMYTCAWRMYTCTCWWYKYTWTGAT
KSTSAACSCKWWGWTCTTAAMMSYKWWKKKYTTWRMAKYKTWTRRYWWKGAAKCRYTWMW
YWASKMWKCGSKAYYTYRYTSWKSKKSWWSKYKTWKAKKYWMTMRKWYWMWWKWCAWYYA

```

It looks like Julia might be stumbling over itself at some point. Maybe I am missing something simple?

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [October 15, 2024, 5:23pm UTC](https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341/2 "2024-10-15T17:23:56Z")

</div>

Is it possible that joe does not search through the whole file but only through some loaded buffer? Have you tried `grep Y yourfile`?

---

<div class="post-metadata">

**Author:** ![Colin\_Beckingham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/colin_beckingham/32/4878_2.png) [@Colin\_Beckingham](https://discourse.julialang.org/u/Colin_Beckingham)\
**Post date:** [October 15, 2024, 5:34pm UTC](https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341/3 "2024-10-15T17:34:26Z")

</div>

Thanks, `grep Y` shows the “bad” sections in the original file.

In the interests of completeness, I should add that in fact the original file appears to be correct. I downloaded the original file a few more times and produced the same answer each time. Y, along with many other characters, are valid in a genome string in addition to the ACGT. See the Wikipedia entry for FASTA file for further information.

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [October 15, 2024, 6:37pm UTC](https://discourse.julialang.org/t/using-readlines-with-large-text-files/121341/4 "2024-10-15T18:37:37Z")

</div>

You might be interested in [GitHub - BioJulia/FASTX.jl: Parse and process FASTA and FASTQ formatted files of biological sequences.](https://github.com/BioJulia/FASTX.jl) and other packages from the BioJulia org
