# Genetic Epidemiology tools

**URL:** <https://discourse.julialang.org/t/genetic-epidemiology-tools/117764>\
**Category:** Biology, Health, and Medicine\
**Created:** [August 2, 2024, 12:32pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764 "2024-08-02T12:32:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![jromanowska](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jromanowska/32/211061_2.png) [@jromanowska](https://discourse.julialang.org/u/jromanowska)\
**Post date:** [August 2, 2024, 12:32pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/1 "2024-08-02T12:32:18Z")

</div>

Hi,  
I am working with genotyping data (PLINK format) and DNA methylation data (large matrices of numbers in range 0-1). Are there any recommended packages, workflows, or tutorials for working with these data in Julia?  
I’m interested in performing GWAS as well as QTL-type analyses.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [August 2, 2024, 3:35pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/2 "2024-08-02T15:35:54Z")

</div>

First of all, welcome!

Second - this was talked about in Slack, but it might be nice to get it here for posterity (slack messages vanish after a couple of weeks).

@Mateusz_K suggested VariantCallFormat.jl for VCF files

You found [https://openmendel.github.io/SnpArrays.jl/](https://openmendel.github.io/SnpArrays.jl/) for PLINK, and there’s also [GitHub - dmbates/BEDFiles.jl: Routines for reading and manipulating GWAS data in .bed files](https://github.com/dmbates/BEDFiles.jl), though that’s much older, and from BioJulia, there’s [GitHub - BioJulia/BED.jl](https://github.com/BioJulia/BED.jl) (I don’t understand the relationship between BED and PLINK, but it seems like they’re related somehow?

---

<div class="post-metadata">

**Author:** ![jromanowska](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jromanowska/32/211061_2.png) [@jromanowska](https://discourse.julialang.org/u/jromanowska)\
**Post date:** [August 5, 2024, 6:56am UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/3 "2024-08-05T06:56:26Z")

</div>

Thank you! I’ll try one of these. I see that BEDFiles.jl and SnpArrays.jl have both much better documentation - BioJulia could improve their package in this sense.

About the methylation data - it’s basically a large matrix (1000x800,000) with numbers in range 0-1. Do you know a fast way to read such data?

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [August 5, 2024, 11:20am UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/4 "2024-08-05T11:20:39Z")

</div>

> [@jromanowska](#):
>
> About the methylation data - it’s basically a large matrix (1000x800,000) with numbers in range 0-1

Are a lot of those numbers 0? If so, you might​want to consider SparseArrays.jl. If not (or in addition), depending on the precision you need, you’ll probably want to use 32 or even 16 bit floats (Julia’s default is 64), to save on memory.

> [@jromanowska](#):
>
> Do you know a fast way to read such data?

Depends on the file type. CSV.jl is very fast, but if it’s a super simple format, readdlm from Base may be enough. @jakobnissen is probably the most knowledgeable in BioJulia about how to go from file bytes to Julia data structures - I’d defer to him.

---

<div class="post-metadata">

**Author:** ![jromanowska](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jromanowska/32/211061_2.png) [@jromanowska](https://discourse.julialang.org/u/jromanowska)\
**Post date:** [August 5, 2024, 12:16pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/5 "2024-08-05T12:16:59Z")

</div>

No, these are not sparse matrices, I can’t use this package. I tried readdlm with Float32, but it’s still quite big. Anyway - will keep trying, thanks for the help so far!

---

<div class="post-metadata">

**Author:** ![jromanowska](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jromanowska/32/211061_2.png) [@jromanowska](https://discourse.julialang.org/u/jromanowska)\
**Post date:** [August 6, 2024, 9:05am UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/6 "2024-08-06T09:05:56Z")

</div>

Just for the record, I’ve found a nice set of tools and tutorials for whole-genome analyses: [JuliaHub](https://juliahub.com/ui/Packages/General/JWAS) 🙂

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [August 6, 2024, 1:13pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/7 "2024-08-06T13:13:59Z")

</div>

For storing & fast reading, I’d give HDF5 a shot: [GitHub - JuliaIO/HDF5.jl: Save and load data in the HDF5 file format from Julia](https://github.com/JuliaIO/HDF5.jl).

---

<div class="post-metadata">

**Author:** ![jromanowska](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jromanowska/32/211061_2.png) [@jromanowska](https://discourse.julialang.org/u/jromanowska)\
**Post date:** [August 6, 2024, 3:59pm UTC](https://discourse.julialang.org/t/genetic-epidemiology-tools/117764/8 "2024-08-06T15:59:24Z")

</div>

Oh, haven’t heard about this one! I’ll definitely give it a try!
