# \[ANN\] Fread.jl - read CSVs faster with the help of R's {data.table}

**URL:** <https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681>\
**Category:** Package Announcements\
**Tags:** performance, data, csv\
**Created:** [October 9, 2019, 1:21pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681 "2019-10-09T13:21:43Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 9, 2019, 1:21pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/1 "2019-10-09T13:21:43Z")

</div>

**Update**  
You should really be using CSV.jl because it performs quite well. I only use Fread.jl for converting data from parquet to feature etc now and not for reading CSVs as CSV.jl is actually really good.

**Original content**

Let’s be honest, all we care about is the speed of CSV reading. I think it was one of @jeff.bezanson’s quotes.

@quinnj has been putting some great to make CSV.jl pretty awesome! It’s getting pretty close to being absolutely awesome! See [Refactor internals to allow better memory efficiency by quinnj · Pull Request #510 · JuliaData/CSV.jl · GitHub](https://github.com/JuliaData/CSV.jl/pull/510)

However, there is no beating of 10 years of fine-tuned awesomeness by the [{data.table}](https://rdatatable.gitlab.io/data.table/) crew!

So I am bring the {data.table} awesomeness to Julia via [Fread.jl](https://github.com/xiaodaigh/Fread.jl)

```julia
using Fread
df = fread("path/to/your/file.csv")

```

want to use {data.table}'s [other arguments](https://rdatatable.gitlab.io/data.table/reference/fread.html#arguments)? You need to use `arg = ` explicitly. E.g.

```julia
using Fread
df = fread("path/to/your/file.csv", sep = "|", nrows = 5000)

```

It should be faster for reading large CSVs than all native Julia CSV reader at the moment (including CSV.jl#jq/mem3 as of 20191010 on Julia 1.3-rc3).

Here are two benchmarks

![txtplot_read](https://global.discourse-cdn.com/julialang/original/3X/1/5/15101afc1d51e639aefe9bb343e7e41f4be159c9.png) ![csvplot_read](https://global.discourse-cdn.com/julialang/original/3X/9/7/97c4861a7e64e84f46a399df4d2e0e0514078d04.png)

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 9, 2019, 2:20pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/2 "2019-10-09T14:20:32Z")

</div>

One more benchmark

![image](https://global.discourse-cdn.com/julialang/original/3X/6/e/6ec572c15cf949a244180271f8b702c78212a638.png)

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [October 9, 2019, 2:54pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/3 "2019-10-09T14:54:13Z")

</div>

Just from a quick browse of the code, it seems like you use R’s `fread` to read the table and write it as a feather file, and then use Julia’s native `Feather` package to read it back in.

Just wondering if this back and forth is still faster than `CSV.jl` and if so, why? What makes the `CSV.jl` package much slower in this case?

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 9, 2019, 4:02pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/4 "2019-10-09T16:02:15Z")

</div>

I think one potential issue here is that `Feather.read` doesn’t actually load the data from disc, it just reads the meta-data, and then the data will get loaded from disc when you actually access values. Not sure how these benchmarks were run, but potentially they didn’t include the reading of the data from disc from the feather files back into memory.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 9, 2019, 9:07pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/5 "2019-10-09T21:07:59Z")

</div>

Oh yes the data is just mapped for feather. I have tried to run a group by after though. The performance is decent.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 9, 2019, 9:14pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/6 "2019-10-09T21:14:19Z")

</div>

Firstly reading from feather doesn’t actually read. It just maps. Secondly, fread is very mature (10 years of development), so it’s much faster. If Julia has arrow then the data can be accessed quicker without going thru the feather step.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 9, 2019, 9:35pm UTC](https://discourse.julialang.org/t/ann-fread-jl-read-csvs-faster-with-the-help-of-rs-data-table/29681/7 "2019-10-09T21:35:31Z")

</div>

If anyone finds issue with the benchmark in anyway feel free

1. Try it on your own data
2. suggest a way to eliminate the issues

Happy to incorporate. I think Fread.jl will die once Julia has a world beating CSV reader which @quinnj is working on.
