# What's the difference between CSV.jl and CSVFiles.jl?

**URL:** <https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950>\
**Category:** New to Julia\
**Created:** [November 24, 2018, 7:51pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950 "2018-11-24T19:51:46Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 24, 2018, 7:51pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/1 "2018-11-24T19:51:46Z")

</div>

What’s the difference between CSV.jl and CSVFiles.jl?  
When should I use each of them?

> **[GitHub - queryverse/CSVFiles.jl: FileIO.jl integration for CSV files](https://github.com/queryverse/CSVFiles.jl)**
>
> FileIO.jl integration for CSV files. Contribute to queryverse/CSVFiles.jl development by creating an account on GitHub.

> **[GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other...](https://github.com/JuliaData/CSV.jl)**
>
> Utility library for working with CSV and other delimited files in the Julia programming language - GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other delimited files in the J...

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [November 24, 2018, 9:15pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/2 "2018-11-24T21:15:01Z")

</div>

They are two independent CSV readers. You should use whichever suits your needs better 🙂

I maintain [CSVFilesjl](https://github.com/queryverse/CSVFiles.jl), so I can probably better speak to that one. Here are some things I like about it:

- It is part of the larger [Queryverse.jl](https://github.com/queryverse/Queryverse.jl) file IO story, which gives you a nice uniform API not just for CSV files, but also [ExcelFiles.jl](https://github.com/queryverse/ExcelFiles.jl), [FeatherFiles.jl](https://github.com/queryverse/FeatherFiles.jl), [StatFiles.jl](https://github.com/queryverse/StatFiles.jl) (Stata, SPSS, and SAS files), and [ParquetFiles.jl](https://github.com/queryverse/ParquetFiles.jl). All of that is documented [here](http://www.david-anthoff.com/jl4ds/stable/fileio.html).
- It works with any source or sink that implements the [TableTraits.jl](https://github.com/queryverse/TableTraits.jl) interface. When I last counted that was something like 21 packages.
- It uses [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) under the hood, which is [fast](https://discourse.julialang.org/t/textparse-jl-is-fast-again/16664). And getting faster, there are a bunch of things on `master` and in branches that are not yet reflected in the benchmarks I posted there.
- It is a mature package, by julia standards. Its been around for about 1.5 years. It gets a continues stream of improvements, but the basic structure has been settled and battle tested since the beginning. It has been a very stable story, and I expect that to stay that way going forward, i.e. I generally try very hard to not break things, and the track record so far has been pretty good, I think 🙂
- You can read gz compressed files out of the box. I’m just mentioning this here, because in a hilarious twist the package has supported that for a long time, but I only found out a few weeks ago 🙂 See [here](https://github.com/queryverse/CSVFiles.jl/issues/33).

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 24, 2018, 11:08pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/3 "2018-11-24T23:08:23Z")

</div>

Thank you.  
Does it deal well with missings?  
Do you suggest using directly CSVFiles.jl or TextParsejl or Queryverse.jl to read and write large csv files?  
And what about binary large binary files?

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [November 24, 2018, 11:40pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/4 "2018-11-24T23:40:34Z")

</div>

Yes, it should deal well with missing values. In fact, the missing handling of the [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl)/[TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) combo is especially robust: even if the type detection algorithm of the CSV parser initially doesn’t recognize that a column can contain missing values, it will still all work if it then comes across a missing value somewhere else in the file.

Right now there is a small overhead of using [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) vs raw [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl), but I think overall it is not too bad, so I would still probably use [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) because it is more convenient. Once the next version of [DataFrames.jl](https://github.com/JuliaData/DataFrames.jl) is released, that overhead will be gone (you can already get that behavior by using the `master` branch of [DataFrames.jl](https://github.com/JuliaData/DataFrames.jl) today).

[Queryverse.jl](https://github.com/queryverse/Queryverse.jl) is just a meta-package that loads ([https://github.com/queryverse/CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl)) and a lot of other packages. If you only want to use ([https://github.com/queryverse/CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl)), I would just load that. But in any case, you will get the same functionality/code, no matter which of these two packages you load.

I like to use [FeatherFiles.jl](https://github.com/queryverse/FeatherFiles.jl) for a binary table format. The same caveat re [DataFrames.jl](https://github.com/JuliaData/DataFrames.jl) applies there as well, you’ll get much better performance with `DataFrames#master` today.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [November 30, 2018, 6:19am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/5 "2018-11-30T06:19:24Z")

</div>

I put together a quick comparison of supported features between CSVFiles.jl & CSV.jl:

| | CSV.jl | CSVFiles.jl |
| --- | --- | --- |
| `Char` delimiters | ✅ | ✅ |
| `String` delimiters | ✅ | |
| Space delimiter | ✅ | ✅ |
| Ignoring any repeated delimiter (fix-width files) | ✅ | |
| `Char` quote characters | ✅ | ✅ |
| Separate open/close quote characters | ✅ | |
| escape characters | ✅ | ✅ |
| skip rows to parse | ✅ | ✅ |
| limit rows to parse | ✅ | |
| handle files without header row | ✅ | ✅ |
| manually provide column names | ✅ | ✅ |
| specify arbitrary row or range of rows for column headers | ✅ | |
| specify # of rows to use for type inference | | ✅ |
| automatically sample entire file for type inference | ✅ | |
| manually specify column types by index or name | ✅ | ✅ |
| parse all columns of one type as another | ✅ | |
| specify arbitrary missing value strings | ✅ | ✅ |
| transform column names into valid julia identifiers | ✅ | |
| specify arbitrary row where “data” begins | ✅ | |
| skip parsing # of rows at end of file | ✅ | |
| specify comment character/string to ignore commented rows | ✅ | |
| read a file transposed (rows as columns, columns as rows) | ✅ | |
| support alternative decimal separator characters (`3.14` vs. `3,14`) | ✅ | |
| specify arbitrary values to parse as `true` or `false` `Bool` values | ✅ | |
| auto detect and parse columns as `CategoricalArray` | ✅ | |
| option to ignore invalid values (replaced with `missing`) | ✅ | |
| ability to selectively parse specific columns | ✅ | |
| ability to apply arbitrary scalar transform functions while parsing | ✅ | |

CSV.jl has been around for a long time now (started May 2015) vs. CSVFiles.jl (June 2017), and developed a ton of requested features. Please chime in if there’s anything missing!

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [November 30, 2018, 7:08pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/6 "2018-11-30T19:08:15Z")

</div>

That is super helpful!

Two additional features that come to mind that [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl) supports is loading data directly from a URL and native support for gz compressed files. There are a couple of WIP PRs over at [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) that will add a few more features from @quinnj’s list to [CSVFiles.jl](https://github.com/queryverse/CSVFiles.jl). Oh, and I guess one other feature of [TextParse.jl](https://github.com/JuliaComputing/TextParse.jl) is the pretty robust column promotion stuff: if your type detection algorithm classifies a column say as integer, but then later in the file you come across a float, it will still just work. Same for missing data: if the type detection algorithm classifies a column to not have missing data, everything still works if later on missing data appears. The promotion is not perfect at this point, i.e. there are some important cases that it doesn’t support, but especially the missing value story is nice and works well, IMO.

---

<div class="post-metadata">

**Author:** ![FredC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredc/32/9857_2.png) [@FredC](https://discourse.julialang.org/u/FredC)\
**Post date:** [October 25, 2019, 8:10pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/7 "2019-10-25T20:10:47Z")

</div>

I did some benchmarks and CSVFiles is an order of magnitude more resource intensive than CSV, and twice as slow even on a second run, i.e. after being compiled.

I’m lumping Queryverse in with the Julia Computing “products” as “considered harmful” packages that should be taken out of the registry.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [October 25, 2019, 8:55pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/8 "2019-10-25T20:55:26Z")

</div>

Your notion of what role Julia Computing plays in the Julia ecosystem seems like it may be misguided. This blog post may help clarify:

> **[The Julia Project and Its Entities](https://julialang.org/blog/2019/02/julia-entities/)**
>
> The Julia Project and Its Entities | There are a number of entities surrounding the Julia programming language. Understandably, many people are not entirely clear on what these groups are and what their relationship to each other is. It’s pretty hard...

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 25, 2019, 9:38pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/9 "2019-10-25T21:38:39Z")

</div>

It is pretty well-established now that CSV.jl is the best CSV reader in Julia.

For Queryverse, I know I wouldn’t use Query.jl because it has poor group\_by performance.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [October 25, 2019, 9:39pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/10 "2019-10-25T21:39:24Z")

</div>

> [@FredC](#):
>
> I’m lumping Queryverse in with …  
> packages that should be taken out of the registry.

No need to be rude.  
Just because something doesn’t work for your use-case doesn’t mean it is some kind of antipattern.

Queryverse doesn’t fit my workflow, but I know many others find it useful.  
And it is much more than CSVFiles.  
You can even use CSV.jl with Queryverse, without issues.

Further, given CSV.jl is according to recent benchmark one of the all time fastest CSV parsers. Getting to be ony 2x slower is solid

And we need competition to drive innovation.

Don’t be rude

---

<div class="post-metadata">

**Author:** ![FredC](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredc/32/9857_2.png) [@FredC](https://discourse.julialang.org/u/FredC)\
**Post date:** [October 26, 2019, 12:21am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/11 "2019-10-26T00:21:18Z")

</div>

I apologize. While I’m not a fan of what I’ve used from Queryverse, it hasn’t been the kind of time-burning bait-and-switch I’ve experienced with the Julia Computing packages.

I’m absolutely baffled at the way things like JuliaDB are being recommended and marketed despite being severely broken. There’s something pyrrhic going on there and the Julia community needs to be vigilant, in my opinion.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 26, 2019, 2:47am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/12 "2019-10-26T02:47:26Z")

</div>

> [@Juan](#):
>
> binary large binary files?

Have u tried JDF.jl?

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [October 26, 2019, 3:45am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/13 "2019-10-26T03:45:09Z")

</div>

> [@FredC](#):
>
> CSVFiles is an order of magnitude more resource intensive than CSV

What resources are you referring to? Main memory, or something else? What OS are you on?

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [October 27, 2019, 2:58pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/14 "2019-10-27T14:58:07Z")

</div>

> [@FredC](#):
>
> I’m absolutely baffled at the way things like JuliaDB are being recommended and marketed despite being severely broken.

If you’re talking about JuliaDB in particular, then the main developer started a PhD and left the project. The other main contributors also contribute a lot to things like Julia itself, so the development of that package has stalled somewhat. That happens occasionally for all open source projects. FWIW it is my impression that the issues experienced by JuliaDB were due to the package trying to ambitiously push some technological boundaries.

In my view JuliaComputing plays a unanimously positive role in furthering julia and supporting the community. If there are particular things you think the community needs to be vigilant about, could you specify it then? That makes it a lot easier to engage with.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [October 27, 2019, 7:10pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/15 "2019-10-27T19:10:59Z")

</div>

What alternative do we have instead of JuliaDB?

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [October 27, 2019, 8:18pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/16 "2019-10-27T20:18:56Z")

</div>

> What alternative do we have instead of JuliaDB?

So so many. Dending what you want  
I think there are like over a dozen packages that support the [Tables.jl](https://github.com/JuliaData/Tables.jl) interface.  
Though many are specific purpose, like for various file loading,  
e.g. Feather.jl, CSV.jl etc  
Some are for databases, like SQLLite.jl, and LibPQ.jl  
And some are general purpose dataframes,  
Like DataFrames.jl, TypedTables.jl, Pandas.jl, StructArrays.jl, and IndexedTables.jl (which is a large part of the core of JuliaDB).  
And that is not to mention the standard row table: Vector or NamedTuples, and standard column table: NamedTuple of Vectors.

Of course it really depends what you are using JuliaDB or something else for.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [October 27, 2019, 8:37pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/17 "2019-10-27T20:37:39Z")

</div>

I meant something ready to work alongside packages such as OnlineStats.jl or able to be used to perform easily statistical ondisk operations, reshape tables, add columns and rows, compute by groups or somethig more complex not fitting on memory.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [October 27, 2019, 9:32pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/18 "2019-10-27T21:32:37Z")

</div>

I suspect you can use CSV.jl in Rows mode for that.  
@quinnj ?

Though you really are describing the things that are more or less JuliaDB’s features.  
So if you haven’t confirmed that it doesn’t work for you then I would definately be trying that first,  
and raising issues at least so he problems can be tracked.

---

<div class="post-metadata">

**Author:** ![simonbyrne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simonbyrne/32/19_2.png) [@simonbyrne](https://discourse.julialang.org/u/simonbyrne)\
**Post date:** [October 28, 2019, 5:16pm UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/19 "2019-10-28T17:16:27Z")

</div>

9 posts were split to a new topic: [Why do you use JuliaDB?](https://discourse.julialang.org/t/why-do-you-use-juliadb/30410)

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [January 28, 2020, 12:24am UTC](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950/20 "2020-01-28T00:24:46Z")

</div>

What package let’s you read a file using a “select=columnnames” option to select the columns you want?

The examples I’ve seen so far do read all columns and makes selection later.

> df = CSV.File(“cool\_file.csv”) |\> select(:a, :b) |\> DataFrame

This doesn’t seem efficient and won’t alloy you to read files larger than memory.

or

> f = CSV.File(file)  
> for row in f  
> println(“a=(row.a), b=(row.b)”)  
> end

Is there an option on any Julia package to read only the desired columns?

[Next page](https://discourse.julialang.org/t/whats-the-difference-between-csv-jl-and-csvfiles-jl/17950.md?page=2)
