# How to read a large TSV file using Julia?

**URL:** <https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300>\
**Category:** General Usage\
**Tags:** question, csv, io\
**Created:** [October 24, 2021, 6:05pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300 "2021-10-24T18:05:25Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![leon](https://avatars.discourse-cdn.com/v4/letter/l/dc4da7/32.png) [@leon](https://discourse.julialang.org/u/leon)\
**Post date:** [October 24, 2021, 6:05pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/1 "2021-10-24T18:05:25Z")

</div>

I have a very large TSV file, about 10GB, it has ~10,000 text lines for metadata information (different versions can have different total number of metadata lines), followed by a row of header information for about 20 parameters, then followed the millions and millions of rows of data.

What is the best way to read such a file in Julia? I tried to use `CSV.read` but it did not work. Any recommendations are greatly appreciated. Many thanks!

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 24, 2021, 6:05pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/2 "2021-10-24T18:05:53Z")

</div>

> [@leon](#):
>
> What is the best way to read such a file in Julia? I tried to use `CSV.read` but it did not work. Any recommendations are greatly appreciated. Many thanks!

In what way did it not work? Did you use the keyword argument `delim = '\t'`?

---

<div class="post-metadata">

**Author:** ![leon](https://avatars.discourse-cdn.com/v4/letter/l/dc4da7/32.png) [@leon](https://discourse.julialang.org/u/leon)\
**Post date:** [October 24, 2021, 6:17pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/3 "2021-10-24T18:17:30Z")

</div>

Many thanks for the tip! It works much better with that keyword argument. However, I’m getting a ton of errors like the below:

```julia
**┌ Warning:** thread = 1 warning: only found 1 / 15 columns around data row: 6618. Filling remaining columns with `missing`
**└** @ CSV ~/.julia/packages/CSV/b4GfC/src/file.jl:622

```

Is there a way to ignore the 10,000 or so lines of metadata information, but only read the header and values into a dataframe?

> [@pdeffebach](#):
>
> delim = ‘\t’

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 24, 2021, 6:20pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/4 "2021-10-24T18:20:48Z")

</div>

Maybe the keyword argument `skipto` will help? See [here](https://csv.juliadata.org/stable/reading.html#skipto).

---

<div class="post-metadata">

**Author:** ![leon](https://avatars.discourse-cdn.com/v4/letter/l/dc4da7/32.png) [@leon](https://discourse.julialang.org/u/leon)\
**Post date:** [October 24, 2021, 6:40pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/5 "2021-10-24T18:40:12Z")

</div>

Many thanks. This is very helpful. I would definitely need this argument in my code.

I wonder if anyone has experiences of detecting the number of lines that should be skipped heuristically. That way, I do not need to worry about crashing my computer by trying to open such a big file using Excel.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 24, 2021, 7:05pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/6 "2021-10-24T19:05:12Z")

</div>

> [@leon](#):
>
> do not need to worry about crashing my computer by trying to open such a big file using Excel.

Consider opening the large text file using an adequate viewer for Windows. One that can be obtained for free from Microsoft store is [Large text viewer](https://www.microsoft.com/en-us/p/large-text-viewer/9nblggh4mcm8#activetab=pivot:overviewtab).

Example of viewing file with 15GB (200M rows x 4 columns):

 ![Windows_Large_text_viewer_df_200M_x_4_tsv](https://global.discourse-cdn.com/julialang/original/3X/0/f/0f260955dfa78b027bd85c6cf822b774e5be7061.png)

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 24, 2021, 10:27pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/7 "2021-10-24T22:27:25Z")

</div>

Try the CSV Chunk reader in DataConvenience.jl if the one in CSV.jl doesn’t suit your needs

[https://github.com/xiaodaigh/DataConvenience.jl#csv-chunk-reader](https://github.com/xiaodaigh/DataConvenience.jl#csv-chunk-reader)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 24, 2021, 11:17pm UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/8 "2021-10-24T23:17:45Z")

</div>

I did something similar (for smaller files however). The files I wanted to import were CSV but there were multiple distinct tables in each file as well as sporadically placed notes. I read the entire file as a string and then had some heuristics for which rows were notes, which were part of a table, and when I had started or ended a table. Then I just pass the table chunks as a String to CSV.jl for parsing. If you know you just have one variable length section of notes followed by only a single table, you could easily do something similar.

---

<div class="post-metadata">

**Author:** ![leon](https://avatars.discourse-cdn.com/v4/letter/l/dc4da7/32.png) [@leon](https://discourse.julialang.org/u/leon)\
**Post date:** [October 25, 2021, 1:29am UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/9 "2021-10-25T01:29:51Z")

</div>

Thanks for sharing.

I’m doing the same thing right now. I read all rows using `readlines`, and then using `occursin` to check where my header row is located. It’s been working well.

---

<div class="post-metadata">

**Author:** ![leon](https://avatars.discourse-cdn.com/v4/letter/l/dc4da7/32.png) [@leon](https://discourse.julialang.org/u/leon)\
**Post date:** [October 25, 2021, 1:32am UTC](https://discourse.julialang.org/t/how-to-read-a-large-tsv-file-using-julia/70300/10 "2021-10-25T01:32:28Z")

</div>

I’m a Mac user and wonder if you know what’s the best Mac alternative? I did some google search and people are talking about BBEdit? I installed Hex Friend, it is truly fast except that it is binary and I can’t really read the information and tell which row contains what.
