# Using bgzipped VCF files with GeneticVariation.jl

**URL:** <https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187>\
**Category:** Biology, Health, and Medicine\
**Tags:** question\
**Created:** [August 2, 2017, 4:25pm UTC](https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187 "2017-08-02T16:25:48Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![weinstockj](https://avatars.discourse-cdn.com/v4/letter/w/b5e925/32.png) [@weinstockj](https://discourse.julialang.org/u/weinstockj)\
**Post date:** [August 2, 2017, 4:25pm UTC](https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187/1 "2017-08-02T16:25:48Z")

</div>

I’m interested in using GeneticVariation.jl to work with VCF files. My VCF files are compressed with bgzip and are indexed using tabix.

I’m looking at the documentation for reading in VCF files here: [https://github.com/BioJulia/GeneticVariation.jl/blob/master/docs/src/io/vcf-bcf.md](https://github.com/BioJulia/GeneticVariation.jl/blob/master/docs/src/io/vcf-bcf.md) . The document suggests to me that the VCF reader should be used with uncompressed VCF’s.

It is my understanding that I could change  
`reader = VCF.Reader(open("example.vcf", "r"))`  
to  
`reader = VCF.Reader(open(`zless example.vcf.gz`, "r"))`

To read in a compressed VCF. Is this the recommended way for working with compressed VCF files? It seems that there’s a reader for BCF files as well, though I’d rather not have to convert all of my vcf.gz files to bcfs.

---

<div class="post-metadata">

**Author:** ![js135005](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/js135005/32/8219_2.png) [@js135005](https://discourse.julialang.org/u/js135005)\
**Post date:** [August 2, 2017, 10:20pm UTC](https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187/2 "2017-08-02T22:20:21Z")

</div>

You might want to look at the Libz.jl package. It can handle .gz files directly and is almost as fast as reading uncompressed files directly.

---

<div class="post-metadata">

**Author:** ![bicycle1885](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bicycle1885/32/107_2.png) [@bicycle1885](https://discourse.julialang.org/u/bicycle1885)\
**Post date:** [August 3, 2017, 1:12am UTC](https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187/3 "2017-08-03T01:12:33Z")

</div>

Hi, I’m a developer of both GeneticVariation.jl and Libz.jl.

I’d rather recommend using [CodecZlib.jl](https://github.com/bicycle1885/CodecZlib.jl) to decompress gzip files. CodecZlib.jl is a member of [TranscodingStreams.jl](https://github.com/bicycle1885/TranscodingStreams.jl), which offers a consistent APIs to various compression formats. I’m going to support automatic gzip decompression using CodecZlib.jl in our bio packages, but now, you can write it like:

```julia
using GeneticVariation
using CodecZlib

reader = VCF.Reader(GzipDecompressionStream(open("example.vcf.gz"))

```

---

<div class="post-metadata">

**Author:** ![weinstockj](https://avatars.discourse-cdn.com/v4/letter/w/b5e925/32.png) [@weinstockj](https://discourse.julialang.org/u/weinstockj)\
**Post date:** [August 3, 2017, 7:46pm UTC](https://discourse.julialang.org/t/using-bgzipped-vcf-files-with-geneticvariation-jl/5187/4 "2017-08-03T19:46:41Z")

</div>

This seemed to do the trick. Thanks for the answer and for developing GeneticVariation.jl + Libz.jl!
