# \[ANN\] New CSV.jl 0.5 Release

**URL:** <https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848>\
**Category:** Package Announcements\
**Tags:** data, csv\
**Created:** [May 4, 2019, 1:53pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848 "2019-05-04T13:53:48Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [May 4, 2019, 1:53pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/1 "2019-05-04T13:53:48Z")

</div>

I’m pleased to announce a new release of the [CSV.jl](https://github.com/JuliaData/CSV.jl) package.

This release provides notable improvements in several areas, including performance, additional features, and enhanced flexibility.

Notable improvements:

- “Perfect” column typing: gone are the days of `rows_for_type_detect` and parsing getting messed up after 10K rows. CSV.jl now gets column types right _ **every time** _, and without needing to restart parsing.
- Auto delimiter detection: don’t worry about keeping track of which file has which delimiter; CSV.jl will figure it out for you!
- Better automatic handling of invalid files: invalid values? wrong number of values on a row? CSV.jl will handle such files gracefully, printing helpful messages about anything unexpected it runs into
- Improved performance: great care has been taken to improve performance on several levels; underlying type parsers (provided by [Parsers.jl](https://github.com/JuliaData/Parsers.jl)), better data locality and cache friendliness, and greater use of custom Julia structures for efficiency
- Enhanced APIs for the [`CSV.File`](http://juliadata.github.io/CSV.jl/dev/#CSV.File) type: in addition to allowing iteration over rows directly, it now provides `getproperty` to access efficient read-only columns of the underlying data. If mutable columns are needed, you can `copy(col)` or use `CSV.read(file; copycols=true)`.
- And lots and lots of [examples](http://juliadata.github.io/CSV.jl/dev/#Examples-1) in documentation!

These improvements are in addition to many smaller bugfixes and quality of life enhancements. Great effort has been taken to ensure CSV.jl provides a rich set of features, comparable or better than other world-class csv parsers. (see the feature comparison table below!)

As always, please open [issues](https://github.com/JuliaData/CSV.jl/issues) as you run into bugs or performance issues and we’ll try to address things as quickly as possible. Cheers!

| | CSV.jl | R fread | Pandas |
| --- | --- | --- | --- |
| `Char` delimiters | ✔ | ✔ | ✔ |
| `String` delimiters | ✔ | | ✔ |
| `Regex` delimiters | | | ✔ |
| Fixed-width files | ✔ | | ✔ |
| Quoted fields | ✔ | ✔ | ✔ |
| Custom open/close quote characters | ✔ | | |
| Skip/offset rows to parse | ✔ | ✔ | ✔ |
| Limit rows to parse | ✔ | ✔ | ✔ |
| Manually provide column names | ✔ | ✔ | ✔ |
| Multiple rows as column names | ✔ | | ✔ |
| Perfect type inference w/o restarting | ✔ | | |
| Manually specify column types | ✔ | ✔ | ✔ |
| Specify arbitrary missing strings | ✔ | ✔ | ✔ |
| “Normalize” column names | ✔ | ✔ | ✔ |
| Skip rows at end of file | ✔ | | ✔ (python engine only) |
| Ignore commented rows | ✔ | | ✔ |
| Handle rows w/ too few/many columns | ✔ | ✔ | ✔ |
| Read a file transposed | ✔ | | |
| Custom decimal separator for floats | ✔ | ✔ | ✔ |
| Custom Bool string values | ✔ | | ✔ |
| “pool” string column values | ✔ | ✔ | |
| Control over invalid values | ✔ | ✔ | ✔ |
| Select/drop specific columns | ✔ | ✔ | ✔ |
| Apply transform functions | ✔ | | ✔ |
| Iteration over rows | ✔ | | ✔ |
| Able to parse Date/DateTime values | ✔ | | ✔ |
| Support reading any IO object | ✔ | ✔ | ✔ |
| Progress meter while parsing | | ✔ | |
| Non-UTF-8 encoded files | | ✔ | ✔ |
| Multi-threaded parsing | | ✔ | |

---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [May 4, 2019, 4:37pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/2 "2019-05-04T16:37:05Z")

</div>

Great set of examples in the docs!

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 4, 2019, 4:54pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/3 "2019-05-04T16:54:02Z")

</div>

> [@quinnj](#):
>
> Improved performance: great care has been taken to improve performance on several levels; underlying type parsers (provided by [Parsers.jl](https://github.com/JuliaData/Parsers.jl)), better data locality and cache friendliness, and greater use of custom Julia structures for efficiency

It would be very useful if you (or someone who participated in this optimization) could write up a short blog post or a summary that links the relevant PRs and discusses what lead to practical improvements. Learning from an actual use case that was optimized by experts is always instructive.

---

<div class="post-metadata">

**Author:** ![glwc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/glwc/32/8317_2.png) [@glwc](https://discourse.julialang.org/u/glwc)\
**Post date:** [May 4, 2019, 5:09pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/4 "2019-05-04T17:09:13Z")

</div>

Great and thank you - this is way faster!

Immediately noticed the following parse behaviour change:

`428.E+03` is default parsed as `String` now, used to be as `Float64` with previous version of `CSV.jl`.

If I try to force with `types=Dict(Symbol(" S-Mises")=>Float64)) ` for the column in question, I get warnings/errors like

> warnings: error parsing Float64 on row = 193959, col = 13: " 428.E+03,", error=INVALID: OK | SENTINEL | DELIMITED | INVALID\_DELIMITER

at all lines that are missing a digit between the `.` and the `E` in scientific notation. If I hand edit the source file column value in question to eg. `428.0E+03` the warning/error goes away.

Is this wanted/expected?

Many thanks,  
GC

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 4, 2019, 5:12pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/5 "2019-05-04T17:12:50Z")

</div>

The key trade-off is if you want to work with DataFrames.jl later if you want to materialize a `DataFrame` using `DataFrame` or `DataFrame!` - or equivalently if you use `copycols=true` vs `copycols=false` keyword argument in `CSV.read` (I am not trying to give a comprehensive answer to what @Tamas_Papp wants as I have not developed the change but gives my major consideration when using current design from DataFrames.jl perspective - maybe @quinnj can comment more).

I will use the following `DataFrame` with `10^7` rows as an example (assuming it is written to disk as testfile.txt).

```julia
using Random, DataFrames

Random.seed!(1234)
df = DataFrame(rand(10^7, 10))
df.g = rand(["a", "b", "c"], 10^7)

```

The difference is loading time between `DataFrame` and `DataFrame!` is roughly equal to:

```julia
df = DataFrame!(CSV.File("testfile.txt"));
julia> @btime df2 = DataFrame($df);
  1.302 s (74 allocations: 801.09 MiB)

```

Now when you do some simple `by` operation you get:

```julia
julia> @btime by($df, :g, c=:x5=>sum);
  278.866 ms (155 allocations: 190.74 MiB)

julia> @btime by($df2, :g, c=:x5=>sum);
  120.093 ms (155 allocations: 190.75 MiB)

```

The short story (from my perspective) is that:

- if you read in the data and do only a few operations on it use `DataFrame!` (i.e. `copycols=false`)
- if you read in the data and to a lot of different operations on it probably using `DataFrame` (i.e. `copycols=true`) will be faster.

Of course another consideration is that `DataFrame!` returns a read-only `DataFrame`, so if you want to mutate vectors you have read in you have to use `DataFrame` constructor.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [May 4, 2019, 5:55pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/6 "2019-05-04T17:55:33Z")

</div>

Thanks for reporting! Would you mind opening an issue on the CSV.jl or Parsers.jl repo? This must have regressed with all the new work that’s gone in, should be a simple fix.

---

<div class="post-metadata">

**Author:** ![glwc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/glwc/32/8317_2.png) [@glwc](https://discourse.julialang.org/u/glwc)\
**Post date:** [May 4, 2019, 6:29pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/7 "2019-05-04T18:29:31Z")

</div>

Issue filed in CSV.jl repo. Cheers,  
GC

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [May 5, 2019, 5:22pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/8 "2019-05-05T17:22:36Z")

</div>

Question: would it be possible to get this to infer?

```julia
file = CSV.File("file.csv")
test(it) = first(it).first_column
@code_warntype test(file)

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 6, 2019, 6:38am UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/9 "2019-05-06T06:38:53Z")

</div>

If the column type is determined from the file contents, no.

Use a [function barrier](https://docs.julialang.org/en/v1/manual/performance-tips/#kernel-functions-1).

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [May 6, 2019, 3:37pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/10 "2019-05-06T15:37:18Z")

</div>

But isn’t the type of each column is already in `file`?

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [May 6, 2019, 3:47pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/11 "2019-05-06T15:47:51Z")

</div>

CSV.jl defines `CSV.getcell(f::CSV.File, T, col, row)` which would be inferrable for individual values. It also doesn’t require iteration. You can get the types for a file by doing `CSV.gettypes(f)`.

---

<div class="post-metadata">

**Author:** ![simeonschaub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simeonschaub/32/216566_2.png) [@simeonschaub](https://discourse.julialang.org/u/simeonschaub)\
**Post date:** [May 22, 2019, 5:04pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/13 "2019-05-22T17:04:50Z")

</div>

@quinnj Great update so far, really enjoying using it! Just one small thing I ran into: When my file ends with a line, containing only one comment, a line of missings gets added at the end.  
MWE:

```julia
shell> cat test.csv
a,b,c
1,2,3
4,5,6
# Comment

julia> using CSV

julia> CSV.read("test.csv", comment="#")
3×3 DataFrames.DataFrame
│ Row │ a │ b │ c │
│ │ Int64⍰ │ Int64⍰ │ Int64⍰ │
├─────┼─────────┼─────────┼─────────┤
│ 1 │ 1 │ 2 │ 3 │
│ 2 │ 4 │ 5 │ 6 │
│ 3 │ missing │ missing │ missing │

```

Is this an issue or an error on my part?

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [May 22, 2019, 5:15pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/14 "2019-05-22T17:15:19Z")

</div>

Ah, that sounds like a bug! And I think I know what the fix should be for it. Hold tight.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [May 22, 2019, 5:21pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/15 "2019-05-22T17:21:57Z")

</div>

Ok, fix is up in a PR here: [https://github.com/JuliaData/CSV.jl/pull/440](https://github.com/JuliaData/CSV.jl/pull/440). Once tests pass, I’ll merge and tag a new patch release.

---

<div class="post-metadata">

**Author:** ![simeonschaub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simeonschaub/32/216566_2.png) [@simeonschaub](https://discourse.julialang.org/u/simeonschaub)\
**Post date:** [May 22, 2019, 5:46pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/16 "2019-05-22T17:46:28Z")

</div>

Whoa, that went quick! Thank you, keep up the great work!

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [August 24, 2019, 2:17pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/17 "2019-08-24T14:17:40Z")

</div>

Just an update here for those interested: CSV.jl now has multithreaded parsing support on current #master branch, based on the new multithreading support in Julia version 1.3. I wrote up a quick blogpost with some benchmarking comparisons vs. R’s `fread` and pandas. I’m happy to report that CSV.jl has built upon previous performance efforts and the multithreaded case is competitive with `fread`. It’s a great time for csv reading in Julia! Check out the blogpost here: [Everyone’s Favorite Blogpost: CSV Benchmarks – Traitement de Données](https://quinnj.home.blog/2019/08/24/everyones-favorite-blogpost-csv-benchmarks/).

(also, it looks like I can’t edit the original feature comparison table in the post above, so you’ll have to visualize for yourself that CSV.jl now has a check in the “Multi-threaded parsing” box 😉 ).

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [August 24, 2019, 6:37pm UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/18 "2019-08-24T18:37:06Z")

</div>

Are there any plans of creating a package similar to data.table? Its syntax and reference semantics are very efficient.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 20, 2019, 2:41am UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/19 "2019-10-20T02:41:34Z")

</div>

> [@Yifan\_Liu](#):
>
> syntax and reference semantics

Which parts do you like? I am a big data.table user.

`df[,.N, by1]` is definitely handy but I prefer `dplyr` for easier to read code sometimes. `DataFramesMeta.jl` is the closest thing. Avoid Query.jl if you want fast group-by performance.

---

<div class="post-metadata">

**Author:** ![anon92994695](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@anon92994695](https://discourse.julialang.org/u/anon92994695)\
**Post date:** [October 20, 2019, 2:44am UTC](https://discourse.julialang.org/t/ann-new-csv-jl-0-5-release/23848/20 "2019-10-20T02:44:32Z")

</div>

So glad you all got the column typing down!

This is huge and makes the library/ecosystem for datascience feel so much stronger. Looking forward to more developments.
