# Announcement: DataFrames Future Plans

**URL:** <https://discourse.julialang.org/t/announcement-dataframes-future-plans/266>\
**Category:** Data\
**Tags:** announcement\
**Created:** [November 12, 2016, 1:45pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266 "2016-11-12T13:45:41Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 12, 2016, 1:45pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/1 "2016-11-12T13:45:41Z")

</div>

**UPDATE** : the plan described below is going to be implemented in a different way from what was initially announced. The DataFrames package will remain the same as it is now: instead, the new framework will be provided by the new [DataTables](https://github.com/JuliaData/DataTables.jl) package. This will allow for a less disruptive migration of existing code relying on DataFrames.

**UPDATE 2** : an updated summary is available [in this post](https://discourse.julialang.org/t/announcement-an-update-on-dataframes-future-plans/4646).

### Towards DataFrames 0.9.0

The DataFrames package and the surrounding ecosystem are currently undergoing a deep refactoring in development branches, based on a framework developed over the last two years. This work aims to dramatically improve performance by replacing the [`DataArray`](https://github.com/JuliaStats/DataArrays.jl) type (and its `NA` value representing missingness) with the new [`Nullable`](http://docs.julialang.org/en/stable/manual/types/#nullable-types-representing-missing-values), [`NullableArray`](https://github.com/JuliaStats/NullableArrays.jl) (see [this blog post](http://julialang.org/blog/2015/10/nullablearrays)) and [`CategoricalArray`](https://github.com/JuliaData/CategoricalArrays.jl) types. Please refer to [this blog post](http://www.johnmyleswhite.com/notebook/2014/11/29/whats-wrong-with-statistics-in-julia/) for an explanation of the limitations of the current design based on `DataArray`. **The new framework is planned to be released as version 0.9.0 in early February 2017**

### New APIs and Compatibility Breaks

Despite our efforts to preserve backward compatibility, this change will likely break some existing workflows. The standard indexing approach (inherited from R) will no longer be the recommended interface. Instead, convenient, flexible and efficient high-level APIs inspired by the [dplyr](https://github.com/hadley/dplyr) R package, by SQL or by LINQ will be preferred. Users are encouraged to experiment with these approaches even with the current stable DataFrames release (0.8.x series), via the [DataFramesMeta](https://github.com/JuliaStats/DataFramesMeta.jl) and [Query](http://www.david-anthoff.com/Query.jl/stable/) packages. Eventually, an API based on the [StructuredQueries](https://github.com/davidagold/StructuredQueries.jl) package (see [this blog post](http://julialang.org/blog/2016/10/StructuredQueries)), which is still in development, will be provided. Among other advantages, these high-level APIs will eventually support different data sources, from in-memory data frames to out-of-core databases, with very little code changes.

The new DataFrames release will require adjustments from all packages depending on DataFrames. Until then, development will continue to happen on the `master` branch of the git repository. In many cases, both the new and the old frameworks can be supported in parallel (by supporting both `DataArray` and `NullableArray`): when possible, package authors are encouraged to start porting as soon as possible. The porting work is tracked [in a GitHub issue](https://github.com/JuliaStats/DataFrames.jl/issues/1092); take inspiration from existing pull requests, and do not hesitate to ask for help there if needed.

Motivated users can also experiment with the development version, though be warned that the user experience can currently be frustrating due to incomplete support for `Nullable` in Julia and in high-level APIs. This issue, known as “lifting” (see [this discussion](https://github.com/JuliaLang/julia/pull/18758) and [this one](https://github.com/JuliaLang/julia/pull/19034), as well as linked pages), still requires fundamental changes. **We expect these to be complete by early January 2017 to allow for a progressive migration** ; users are not advised to upgrade to the development version for actual work until then.

### More Changes

The above changes will be coordinated with a related refactoring of the DataFrames codebase to increase modularity and :

- CSV reading and writing support (`readtable` and `writetable`) will be deprecated in favor of the [CSV](https://github.com/JuliaData/CSV.jl) package. Data importation and exportation should more generally be done via the [DataStreams](https://github.com/JuliaData/DataStreams.jl) package (see [this blog post](http://julialang.org/blog/2015/10/datastreams)).
- Functions translating model formulas into model matrices will be moved to a separate [StatsModels](https://github.com/JuliaStats/StatsModels.jl) package, with the goal of eventually supporting any kind of `AbstractTable` (including `DataFrame`), and will also include model-related functions currently in StatsBase. Though this will not happen in the first release, in the end modeling packages should only need to depend on that package, and no longer on DataFrames.
- A new `AbstractTable` interface will be progressively developed in [the eponymous package](https://github.com/davidagold/AbstractTables.jl) to allow writing generic code supporting any kind of tabular data, including `DataFrame`, without depending on the DataFrames package.
- Packages strongly tied to DataFrames (including that package itself) will be moved to the [JuliaData](https://github.com/JuliaData/) organization to keep JuliaStats focused on actual statistics.

We are aware that the transition will certainly be disruptive for users. But we are confident the advantages of the new framework will greatly offset its costs, following state-of-the-art designs like R’s [dplyr](https://github.com/hadley/dplyr) and Python’s [Pandas 2.0](https://pandas-dev.github.io/pandas2/), and taking full advantage of Julia’s flexibility and performance. Your help is welcome to push forward with this roadmap!

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 12, 2016, 1:46pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/2 "2016-11-12T13:46:42Z")

</div>



---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [November 12, 2016, 9:19pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/3 "2016-11-12T21:19:11Z")

</div>

Perhaps you could give it a different name. Then people could transition in their own time…

---

<div class="post-metadata">

**Author:** ![ssfrr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ssfrr/32/3736_2.png) [@ssfrr](https://discourse.julialang.org/u/ssfrr)\
**Post date:** [November 12, 2016, 9:29pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/4 "2016-11-12T21:29:41Z")

</div>

Fortunately it’s possible to stick with the older version of a package if you need to. For package developers just put an upper-bound on `DataFrames` in your `REQUIRE` file. As a user you can use `Pkg.pin` to prevent Julia from updating the package until you’re ready.

---

<div class="post-metadata">

**Author:** ![kevin.squire](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevin.squire/32/62_2.png) [@kevin.squire](https://discourse.julialang.org/u/kevin.squire)\
**Post date:** [November 19, 2016, 5:29pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/5 "2016-11-19T17:29:38Z")

</div>

I actually agree with @cormullion that a new package name might be (or have been?) better. Perhaps even DataFrames2? 😉

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 19, 2016, 6:00pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/6 "2016-11-19T18:00:44Z")

</div>

I don’t know. `DataFrames` is a good name for this. People know the name and will search for it. Deprecating the name in order to help people with an API change pre-1.0 seems like prematurely trying to enforce stability.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 19, 2016, 7:23pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/7 "2016-11-19T19:23:52Z")

</div>

Using a different package name wouldn’t help much, since DataFrames is a dependency of many other packages which will need to choose which version to use (until we have optional dependencies at least). It would be really confusing to have a `DataFrame` object which isn’t supported by e.g. Gadfly because it expects a `DataFrame` type from the other DataFrames package.

On the contrary, pinning the package to the 0.8 version will ensure all dependencies are compatible (if we add upper bounds correctly) until you are ready to make the switch.

---

<div class="post-metadata">

**Author:** ![kevin.squire](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevin.squire/32/62_2.png) [@kevin.squire](https://discourse.julialang.org/u/kevin.squire)\
**Post date:** [November 20, 2016, 6:29am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/8 "2016-11-20T06:29:52Z")

</div>

Okay, fair enough. Cheers!

---

<div class="post-metadata">

**Author:** ![viralbshah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/viralbshah/32/54_2.png) [@viralbshah](https://discourse.julialang.org/u/viralbshah)\
**Post date:** [November 20, 2016, 9:03pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/9 "2016-11-20T21:03:55Z")

</div>

I think this is worth posting as a blog post on the Julia blog.

---

<div class="post-metadata">

**Author:** ![cstjean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cstjean/32/1444_2.png) [@cstjean](https://discourse.julialang.org/u/cstjean)\
**Post date:** [November 21, 2016, 12:00am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/10 "2016-11-21T00:00:58Z")

</div>

> [@nalimilan](#):
>
> The standard indexing approach (inherited from R) will no longer be the recommended interface.

Not a R user — what is this referring to? `df[df[:something] .== 5]` ?

---

<div class="post-metadata">

**Author:** ![jpfairbanks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpfairbanks/32/4500_2.png) [@jpfairbanks](https://discourse.julialang.org/u/jpfairbanks)\
**Post date:** [November 21, 2016, 12:33am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/11 "2016-11-21T00:33:11Z")

</div>

Great news,

I am glad to see progress here, and that it is being communicated like this. Is there a list of issues that need to get done before the release?

I found [Milestone 0.9.0](https://github.com/JuliaStats/DataFrames.jl/milestone/2) and [Issue #1092](https://github.com/JuliaStats/DataFrames.jl/issues/1092). Are these the complete lists? Maybe this is a good use for the new Github Projects feature.

-James

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 21, 2016, 9:32am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/12 "2016-11-21T09:32:06Z")

</div>

> [@cstjean](#):
>
> Not a R user — what is this referring to? df[df[:something] .== 5] ?

Not the best place to ask, please look at the docs or start a new thread.[quote=“jpfairbanks, post:11, topic:266”]  
I found Milestone 0.9.0 and Issue #1092. Are these the complete lists? Maybe this is a good use for the new Github Projects feature.  
[/quote]

Yes, that’s more or less complete. I don’t think we need to use the Projects feature, as most of the work needs to happen outside of DataFrames now (in particular, in NullableArrays, StructuredQueries and Julia Base).

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 23, 2016, 9:41am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/13 "2016-11-23T09:41:15Z")

</div>

> [@viralbshah](#):
>
> I think this is worth posting as a blog post on the Julia blog.

I’m not sure, as blog posts usually have announced work when it was more or less completed. I’d prefer to wait until we actually release the new framework, which is when we should draw the most attention to it.

---

<div class="post-metadata">

**Author:** ![zhangliye](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zhangliye/32/3208_2.png) [@zhangliye](https://discourse.julialang.org/u/zhangliye)\
**Post date:** [February 23, 2017, 2:36am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/14 "2017-02-23T02:36:19Z")

</div>

What are the advantages of DataTables.jl compared with DataFrames.jl? From the perspective of performance, which package should I choose? Thanks!

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 23, 2017, 9:53am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/15 "2017-02-23T09:53:58Z")

</div>

The main advantage of DataTables is type stability (at the column level), but whether it will make your code faster will depend on many things, so it’s hard to tell without trying. Also DataTables hasn’t been really optimized yet, though we’re working on it (e.g. [this PR](https://github.com/JuliaData/DataTables.jl/pull/3)).

If you use high-level APIs like Query.jl it should be easy to switch from one framework to another to compare them. If you work directly with column vectors, I strongly recommend using Julia 0.6, where missing values lifting is supported with element-wise operators like `.+`.

---

<div class="post-metadata">

**Author:** ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Post date:** [February 23, 2017, 11:18am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/16 "2017-02-23T11:18:32Z")

</div>

Where is the documentation for DataTables (the links at [https://github.com/JuliaData/DataTables.jl](https://github.com/JuliaData/DataTables.jl) result in a 404 (page not found) error).

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [February 23, 2017, 11:32am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/17 "2017-02-23T11:32:06Z")

</div>

I wouldn’t just dismiss the smart indexing approach for querying a data table. I think there’s merit in supporting both the indexing & the LINQ/dplyr approach. In R when working with in-memory tables I prefer the indexing approach of data.table both for speed & conciseness. See [r - data.table vs dplyr: can one do something well the other can't or does poorly? - Stack Overflow](http://stackoverflow.com/questions/21435339/data-table-vs-dplyr-can-one-do-something-well-the-other-cant-or-does-poorly) for a discussion, especially the second reply in favor of data.table.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 23, 2017, 1:10pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/18 "2017-02-23T13:10:08Z")

</div>

@sylvaticus The manual didn’t build correctly until recently because Query.jl needed to be updated. Now there remains a small bug that [this PR](https://github.com/JuliaData/DataTables.jl/pull/21) should fix.

@Steven_Sagaert I didn’t say we dismissed that syntax, but currently it requires some changes when moving between DataFrames and DataTables, so that’s not the best way to compare them. Also the syntax to work with `Nullable` isn’t stabilized yet, so it’s easier to work with high-level APIs for now. Finally, performance shouldn’t be higher with direct indexing, at least with optimized querying frameworks. (BTW, DataFrames/DataTable does not support data.table-like advanced syntax, only the basic indexing API.)

---

<div class="post-metadata">

**Author:** ![zhangliye](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zhangliye/32/3208_2.png) [@zhangliye](https://discourse.julialang.org/u/zhangliye)\
**Post date:** [February 23, 2017, 5:49pm UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/19 "2017-02-23T17:49:25Z")

</div>

You can get some information from the man directory.  
[https://github.com/JuliaData/DataTables.jl/blob/master/docs/src/man/getting\_started.md](https://github.com/JuliaData/DataTables.jl/blob/master/docs/src/man/getting_started.md)

---

<div class="post-metadata">

**Author:** ![tshort](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tshort/32/43_2.png) [@tshort](https://discourse.julialang.org/u/tshort)\
**Post date:** [February 24, 2017, 1:52am UTC](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266/20 "2017-02-24T01:52:43Z")

</div>

Another package to consider is [IndexedTables](https://github.com/JuliaComputing/IndexedTables.jl), particularly if you like R’s data.table (I do!).

It’s type stable, and the indexing means lookups and joins should be fast.

Note that this package is not registered, yet.

[Next page](https://discourse.julialang.org/t/announcement-dataframes-future-plans/266.md?page=2)
