# TextParse.jl is fast again

**URL:** <https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664>\
**Category:** Data\
**Tags:** announcement\
**Created:** [October 23, 2018, 3:23am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664 "2018-10-23T03:23:01Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 23, 2018, 3:23am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/1 "2018-10-23T03:23:01Z")

</div>

Some of you might have noticed that [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) (and [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl), which is a small wrapper around [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl)) saw some [major performance regressions](https://github.com/JuliaComputing/TextParse.jl/issues/73) initially on julia 1.0.

I just fixed these and now both packages are back to the kind of performance that we saw on julia 0.6 for them (which was pretty good). If you had given up on either package since moving to julia 1.0, I encourage you to give them another try, they should be very usable again.

As part of that work I also wrote a [benchmark](https://github.com/davidanthoff/csv-comparison) that compares various CSV reading packages on julia, Python and R. The high-level summary is that R’s fread beats pretty much everything, but other than that [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) is (and was on julia 0.6) looking pretty good.

Here are the detailed results:

 ![withoutna](https://global.discourse-cdn.com/julialang/original/3X/a/c/aca9052d8b2d0605c030256d89d990dd10cf4621.png)

[direct link to figure](https://discourse.julialang.org/uploads/short-url/dIJKC30hLSV5sowqrnfrkjEX9SH.png)

 ![withna](https://global.discourse-cdn.com/julialang/original/3X/b/0/b006d7f804ab7c980d0fb415a73c35f0f6bccf29.png)

[direct link to figure](https://discourse.julialang.org/uploads/short-url/vT2fkzrQClyCMCWym5tQaYkXlY2.png)

I used the currently latest tagged version of all packages that are tested. A “; 0.6” in the package name means this run was done on julia 0.6. I ran every benchmark five times. The bar shows the best of those five runs, and all five runs are shown as ticks. The different files have different types of data in the columns: the files with “mixed” in the name have one column of float, int, string, categorical string (a string column that only ever has a few different values) and datetime. The files with “uniform” in the name have 20 columns all with the same data type. “short” in the filename signals that floating point numbers don’t have more than 6 digits.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 23, 2018, 5:46am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/2 "2018-10-23T05:46:21Z")

</div>

A side note is that fread is very fast at most everything that it does!

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [October 23, 2018, 7:40am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/3 "2018-10-23T07:40:41Z")

</div>

Thanks a lot for your effort and these extensive benchmarks. I was wondering about the number of columns per data set, it would be great to see how each approach scales as a function of that. Last time I checked, most julia solutions (except DataFrame’s deprecated `readtable` and Base’s `readdlm`) were struggling with high-dimensional data sets (\>1k columns).

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 23, 2018, 8:26am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/4 "2018-10-23T08:26:31Z")

</div>

I have asked before but I wonder whether Julia van eventually match fread.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 23, 2018, 8:46am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/5 "2018-10-23T08:46:07Z")

</div>

There is no theoretical reason it couldn’t, someone just needs to put in the micro-optimization work.

I think that fixing major performance regressions is important, so I am happy that TextParse.jl is competitive again, but I am not sure that it is super-important to beat `data.table::fread` once we are within 5x. Data reading is usually the least expensive part of a nontrivial computation, and one would not read all the data for large files (when this matters the most) into memory anyway.

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [October 23, 2018, 9:04am UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/6 "2018-10-23T09:04:39Z")

</div>

> [@Tamas\_Papp](#):
>
> large files (when this matters the most)

And to add to that, TextParse.jl looks already quite comparable to fread for larger numbers of rows. If one has to read many small csv files (which is also an important use case), TextParse.jl may be slower for now, but coincidentally this is where CSV.jl seems to do very well.

---

<div class="post-metadata">

**Author:** ![matthieu](https://avatars.discourse-cdn.com/v4/letter/m/da6949/32.png) [@matthieu](https://discourse.julialang.org/u/matthieu)\
**Post date:** [October 23, 2018, 2:23pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/7 "2018-10-23T14:23:58Z")

</div>

Looks great! Out of curiosity, can you expand a bit on why you perform better than CSV for large datasets?

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 23, 2018, 4:13pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/8 "2018-10-23T16:13:39Z")

</div>

> [@jtackm](#):
>
> I was wondering about the number of columns per data set, it would be great to see how each approach scales as a function of that.

Agreed, a PR that adds runs with say 100, 1k and 10k columns would be great!

> [@Tamas\_Papp](#):
>
> I think that fixing major performance regressions is important, so I am happy that TextParse.jl is competitive again, but I am not sure that it is super-important to beat `data.table::fread` once we are within 5x.

I entirely agree! My goal with the recent work was not to write the fastest CSV parser, I mainly just wanted to get TextParse.jl back into a usable state, so I did some very targeted optimizations so that it got back to the excellent performance it had on julia 0.6 (thanks to @shashi!). There are actually a lot of places where one could do more, but for now I thought we should release a version that is back to the old performance.

> [@matthieu](#):
>
> Out of curiosity, can you expand a bit on why you perform better than CSV for large datasets?

No idea 🙂 TextParse.jl was just always really fast, I think @shashi just did some awesome work with it. Also keep in mind that CSV.jl (the current version) is really a young package. Even though the package itself has been around for a long time, I believe the current version is essentially a recent complete rewrite. I would guess that this is simply a case where it takes time to mature a package.

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [October 23, 2018, 10:10pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/9 "2018-10-23T22:10:54Z")

</div>

> [@davidanthoff](#):
>
> Agreed, a PR that adds runs with say 100, 1k and 10k columns would be great!

Sure, I’ll have a look

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 29, 2018, 4:11pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/10 "2018-10-29T16:11:45Z")

</div>

Thanks to a PR from @jtackm we now also have results for files with 200 columns:

 ![cols_200_withoutna](https://global.discourse-cdn.com/julialang/original/3X/1/8/182bb695b7b603ca989f94177179db42c2a31e03.png)

[direct link to figure](https://discourse.julialang.org/uploads/short-url/xz52rXjtUNXtEjY1RHmDtm7FWs7.png)

 ![cols_200_withna](https://global.discourse-cdn.com/julialang/original/3X/5/9/59a0e7d6c8d219825006a53171709949063336d6.png)

[direct link to figure](https://discourse.julialang.org/uploads/short-url/aPdKUjen52oVP3IMR44hWE6ETwH.png)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 29, 2018, 4:27pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/11 "2018-10-29T16:27:01Z")

</div>

Do you have any ideas as to why there seems to be substantial overhead between CSVFiles and TextParse? Sometimes the time difference looks negligible, but typically I see CSVFiles lagging behind, even though it is TextParse under the hood.

---

<div class="post-metadata">

**Author:** ![jtackm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jtackm/32/4784_2.png) [@jtackm](https://discourse.julialang.org/u/jtackm)\
**Post date:** [October 29, 2018, 6:24pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/12 "2018-10-29T18:24:14Z")

</div>

I’m curious as well, I saw in one additional high-dimensional test (5K rows, 6K cols) that differences are even more pronounced (TextParse \< 5s, pandas \< 10s, CSVFiles \> 700s). But since TextParse is already there, I think that should be fixable.

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 29, 2018, 6:40pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/13 "2018-10-29T18:40:27Z")

</div>

> [@tbeason](#):
>
> Do you have any ideas as to why there seems to be substantial overhead between CSVFiles and TextParse?

Yes, I do 🙂 I’ve had some PRs lingering around that address that for many, many months. Most of them are merged now. The only thing left to do is to merge [this](https://github.com/JuliaData/DataFrames.jl/pull/1579), and then there is no real overhead left from using [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl), i.e. you get the same performance that raw [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) gives you.

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 29, 2018, 6:41pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/14 "2018-10-29T18:41:31Z")

</div>

> [@jtackm](#):
>
> I’m curious as well, I saw in one additional high-dimensional test (5K rows, 6K cols) that differences are even more pronounced (TextParse \< 5s, pandas \< 10s, CSVFiles \> 700s).

That is actually _very_ good news, that TextParse seems to perform very well even with a couple thousand of columns!

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [October 30, 2018, 8:38pm UTC](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664/15 "2018-10-30T20:38:52Z")

</div>

7 posts were split to a new topic: [Tables.jl vs TableTraits.jl (was TextParse.jl is fast again)](https://discourse.julialang.org/t/tables-jl-vs-tabletraits-jl-was-textparse-jl-is-fast-again/16982)
