# Julia for processing Next-Generation Sequencing (NGS) datasets

**URL:** <https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371>\
**Category:** Biology, Health, and Medicine\
**Tags:** question, package\
**Created:** [April 23, 2024, 5:35am UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371 "2024-04-23T05:35:44Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bonjour-Lemonde](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bonjour-lemonde/32/24199_2.png) [@Bonjour-Lemonde](https://discourse.julialang.org/u/Bonjour-Lemonde)\
**Post date:** [April 23, 2024, 5:35am UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/1 "2024-04-23T05:35:44Z")

</div>

Hi all! I want to inquire how to use Julia for processing Next-Generation Sequencing (NGS) datasets, especially to merge paired-end sequencing reads. I think currently there seems no suitable packages in Julia as pandaseq in python for doing this. Importantly, quality information in the Illumina reads is important for score and evaluate alignment of paired-end sequences[1](https://doi.org/10.1186/1471-2105-13-31). But it seems no packages in Julia has considered and used this information. Thus I want to know whether there are Julia packages that can handle the assembly problem and also I wanted to know suggestions to use python in Julia as well. THANKS!

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [April 23, 2024, 6:51am UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/2 "2024-04-23T06:51:14Z")

</div>

Welcome, @Bonjour-Lemonde !

That’s right, there is no package in Julia to merge paired-end NGS sequences. Julia currently has a bunch of low-level packages for Bioinformatics, such as FASTX.jl to parse FASTQ files, and BioAlignments.jl to do S/W alignment.

However, basic NGS tasks like read trimming and merging and assembly is usually best done with existing command-line tools which tend to be written in C or C++. For Illumina reads, I would recommend `fastp` for trimming and merging, and `SPAdes` for assembly.

Julia is suitable when you need to do a truly custom analysis, e.g. when developing new techniques in the field. For most standard analyses, I would use existing tools.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [April 23, 2024, 10:58am UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/3 "2024-04-23T10:58:48Z")

</div>

> [@jakobnissen](#):
>
> basic NGS tasks like read trimming and merging and assembly is usually best done with existing command-line tools

Agree with this completely. There’s no reason _in principle_ that Julia couldn’t be used to write such tools, but

1. Given limited resources, no one has considered it worth it to duplicate the effort
2. Julia is not (yet?) a great choice for developing command line tools, which most biologists expect.

Depending on your application, there are some Julia packages for downstream analysis (eg SingleCellProjections.jl if you’re doing scRNAseq), and lots of stuff in the stats/ML space.

---

<div class="post-metadata">

**Author:** ![jonathanBieler](https://avatars.discourse-cdn.com/v4/letter/j/82dd89/32.png) [@jonathanBieler](https://discourse.julialang.org/u/jonathanBieler)\
**Post date:** [April 23, 2024, 1:08pm UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/4 "2024-04-23T13:08:52Z")

</div>

I’m using Julia regularly to do custom read trimming & processing, I think doing something like PANDAseq should be relatively straightforward to implement around existing packages, although that’s maybe more a package developer project than end-user one.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [April 23, 2024, 3:04pm UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/5 "2024-04-23T15:04:54Z")

</div>

> [@jonathanBieler](#):
>
> I’m using Julia regularly to do custom read trimming & processing

Oh neat - any chance you’d be willing to write up a tutorial or cookbook recipe for BioTutorials?

---

<div class="post-metadata">

**Author:** ![jonathanBieler](https://avatars.discourse-cdn.com/v4/letter/j/82dd89/32.png) [@jonathanBieler](https://discourse.julialang.org/u/jonathanBieler)\
**Post date:** [April 23, 2024, 5:24pm UTC](https://discourse.julialang.org/t/julia-for-processing-next-generation-sequencing-ngs-datasets/113371/6 "2024-04-23T17:24:40Z")

</div>

I could, although the difficulty is to find a realistic use case that isn’t too boring (otherwise it’s just [this](https://jonathanbieler.github.io/BioRecordsProcessing.jl/dev/examples/#Transforming-a-FASTA-file)) and public data that goes with it.
