# What's the difference between CSV.jl and CSVFiles.jl?

**URL:** <https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950>\
**Category:** New to Julia\
**Created:** [November 24, 2018, 7:51pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950 "2018-11-24T19:51:46Z")\
**Posts on this page:** 6\
**Page:** 2

<div class="post-metadata">

**Author:** ![BMval](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bmval/32/7647_2.png) [@BMval](https://discourse.julialang.org/u/BMval)\
**Post date:** [January 28, 2020, 12:57am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/21 "2020-01-28T00:57:45Z")

</div>

Thank you.  
I was surprised, CSVFiles is 4.5 times fasted for 1.2Gb files (164000\*2500 dataframe).

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 28, 2020, 1:08am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/22 "2020-01-28T01:08:50Z")

</div>

> [@Juan](#):
>
> What package let’s you read a file using a “select=columnnames” option to select the columns you want?

I added something like that to [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) recently:

```julia
load("foo.csv", colparsers=Dict(:colA=>nothing, :colC=>nothing)) |> DataFrame

```

Essentially when you assign `nothing` as the colparser for a given column, it will be skipped entirely.

What I don’t have yet is a nice (positive) column selection API. My goal is to make

```julia
load("foo.csv") |> @select(:colA, :colB)

```

work with this, i.e. even though it would look as if you are selecting columns after they are read, the design of [Query.jl](https://github.com/queryverse/Query.jl) and [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) is such that I can get this to never actually read any column other than `colA` and `colB`. The goal is to support the full column selection story from the [`@select`](https://www.queryverse.org/Query.jl/stable/standalonequerycommands/#The-@select-command-1) command.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [January 28, 2020, 7:58am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/23 "2020-01-28T07:58:41Z")

</div>

Sorry to interfere, but is it possible to have the same functionality for rows as well? Like ignore rows which fails some tests at the read time?  
May be something like this

```julia
load("foo.csv", rowparsers=Dict(:colA => x -> x > 0)) |> DataFrame

```

One obvious applications is csv with comment lines, another is that loading only subset of data may be more memory efficient.

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 28, 2020, 11:16pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/24 "2020-01-28T23:16:21Z")

</div>

For comments, you can already skip reading them via the `commentchar` keyword (and I just realised that wasn’t documented, I just updated the [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) README).

My plan for arbitrary row filtering is two-fold: I hope to add a true streaming mode of operation to [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl). At that point you could just write:

```julia
load("foo.csv", hypothetical_streaming_flag=true) |>
  @filter(_.colA ==3) |>
  DataFrame

```

and then that would stream things row-by-row, applying the filter per row before things get materialised into the `DataFrame.

---

<div class="post-metadata">

**Author:** ![BMval](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bmval/32/7647_2.png) [@BMval](https://discourse.julialang.org/u/BMval)\
**Post date:** [January 29, 2020, 12:56am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/25 "2020-01-29T00:56:26Z")

</div>

But Unfortunately, CSVFiles.jl write 4 time slower, =((

So, to read use CSVFiles.jl, to write use CSV.jl

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 29, 2020, 1:35am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/26 "2020-01-29T01:35:43Z")

</div>

> [@BMval](#):
>
> But Unfortunately, CSVFiles.jl write 4 time slower

Not sure what is going on there, I’ll take a look. At some point they were pretty similar. Right now CSV.jl is much faster on writing numbers because it uses a better floating point write algorithm (I think), but I think the example at the top shouldn’t be affected by that… Not sure.

[Previous page](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950.md?page=1)
