# Large FASTA datasets?

**URL:** <https://discourse.julialang.org/t/large-fasta-datasets/90786>\
**Category:** Biology, Health, and Medicine\
**Created:** [November 25, 2022, 2:35am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786 "2022-11-25T02:35:33Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![M-PERSIC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m-persic/32/51968_2.png) [@M-PERSIC](https://discourse.julialang.org/u/M-PERSIC)\
**Post date:** [November 25, 2022, 2:35am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/1 "2022-11-25T02:35:33Z")

</div>

Hello! Apologies if this is the wrong place to post this question.

I’m playing around with TranscodingStreams and I want to compare the compression of FASTA files with different codecs. I can make my own randomized FASTA files using BioSequences and FASTX easily enough, but I’m also looking for some real world examples.

I’m looking for a large public FASTA file dataset (at least 100+ files available) that’s well reviewed and varied. If it includes both DNA and RNA FASTA files (or there’s a separate dataset you know of) that would work quite well. If anyone can point me to where such datasets are located that would be awesome!

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [November 25, 2022, 8:02am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/2 "2022-11-25T08:02:43Z")

</div>

FormatSpecimens.jl includes a varied set of FASTA files, but they are small.  
Normally, most large datasets are produced in a homogenous manner, and so do not have much variation of the data inside the dataset.

You’d probably want to make it yourself. I would include the following different kinds of dataset

- Plant genome, especially maize, strawberry or pine, which are (in)famous for their repeats
- Oher eukaryote genome, maybe human or yeast
- Assembly of a metagenomic dataset
- Set of variants of e.g. a virus - these are 99% identical so should have unique compression abilities. Look at a covid or flu DB
- Maybe something from mass spec? I’m not too into it, but if they represent the protein fragments as AA, that would be interesting as well

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [November 25, 2022, 8:31am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/3 "2022-11-25T08:31:30Z")

</div>

Also consider adding amplicon data if it fits your use case (quite different characteristics from Whole Genome Shotgun). Microbiome research produced \>1mio such samples these days, uploaded typically to NCBI SRA or ENA.

---

<div class="post-metadata">

**Author:** ![M-PERSIC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m-persic/32/51968_2.png) [@M-PERSIC](https://discourse.julialang.org/u/M-PERSIC)\
**Post date:** [November 25, 2022, 8:56am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/4 "2022-11-25T08:56:36Z")

</div>

@jakobnissen@jtackm Thank you both for your help! While I was doing some research I actually found some references to an [NCBI FTP server](https://ftp.ncbi.nih.gov/genomes/HUMAN_MICROBIOM/Bacteria/) containing almost 900 MB of \*.fna file data! Far more than I need, but will definitely be creating some artifacts 🙂

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [November 25, 2022, 9:02am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/5 "2022-11-25T09:02:46Z")

</div>

These look to be assemblies of microbes from the human gut. Note that these will have completely different compression characteristics from e.g. genomic repeats or a selection of variants of the same sequence.

---

<div class="post-metadata">

**Author:** ![M-PERSIC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m-persic/32/51968_2.png) [@M-PERSIC](https://discourse.julialang.org/u/M-PERSIC)\
**Post date:** [November 25, 2022, 9:12am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/6 "2022-11-25T09:12:59Z")

</div>

Hm, that is something I should mention in the Discussion section for my paper. As this is for a small research project at uni, I think FormatSpecimens.jl might actually work well enough to try out Zstd dictionary compression. Either or should work fine for the purposes of my paper!

---

<div class="post-metadata">

**Author:** ![fcriscuo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fcriscuo/32/22080_2.png) [@fcriscuo](https://discourse.julialang.org/u/fcriscuo)\
**Post date:** [December 11, 2022, 8:44pm UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/7 "2022-12-11T20:44:33Z")

</div>

I just saw your request, but if you are still looking for real sequences in FASTA format you can download a file of 56K+ DNA sequences from the Sanger Lab’s Catalog of Somatic Mutations in Cancer (COSMIC) database ([Download Files](https://cancer.sanger.ac.uk/cosmic/download)). You’ll have to register for a free account. FASTA data for RNA sequences are typically reverse-transcribed to DNA. Also, sequences from the minus strand are usually reoriented to the 5` to 3` direction of the positive strand ([Negative Strand Coordinates in Fasta Files?](https://www.biostars.org/p/158551/)). Hope this helps.

---

<div class="post-metadata">

**Author:** ![M-PERSIC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/m-persic/32/51968_2.png) [@M-PERSIC](https://discourse.julialang.org/u/M-PERSIC)\
**Post date:** [December 15, 2022, 6:21am UTC](https://discourse.julialang.org/t/large-fasta-datasets/90786/8 "2022-12-15T06:21:59Z")

</div>

Thank you! Will keep it bookmarked for later use.
