# \[ANN\] New and Improved JuliaDB

**URL:** https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719
**Category:** Community
**Tags:** package, announcement
**Created:** [January 31, 2018, 5:37pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719 "2018-01-31T17:37:34Z")
**Posts on this page:** 15
**Page:** 1

<div class="post-metadata">

### Author: ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)
#### Post date: [January 31, 2018, 5:37pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/1 "2018-01-31T17:37:34Z")

</div>

[**JuliaDB**](https://github.com/JuliaComputing/JuliaDB.jl) has had a series of big improvements over the past months. In addition to better performance and an [API refresh](http://juliadb.org/latest/api/), there are a few particular features that we would like to bring to your attention:

1. JuliaDB is now [**closely integrated with OnlineStats**](http://juliadb.org/latest/manual/onlinestats.html). The algorithms in OnlineStats naturally lend themselves to working with very large distributed data sets. As a result, you can now calculate descriptive statistics and perform online statistical learning on distributed data sets within JuliaDB quickly and easily.

2. JuliaDB now supports [**machine learning workflows**](http://juliadb.org/latest/manual/ml.html) with helpful utility functions that extract feature matrices out of raw input data. JuliaDB automatically detects continuous/categorical variables to create one-hot representations and standardized data. This allows immediate use of JuliaDB tables in machine learning algorithms with no additional data wrangling.

3. JuliaDB leverages OnlineStats to [**visualize datasets of unlimited size**](http://joshday.github.io/OnlineStats.jl/latest/visualizations.html#Partitions-1) using a broad selection of descriptive statistics, computed with single-pass distributed algorithms. The system builds fixed-size summaries of infinite data streams as the data comes in.

 ![randwalk](https://global.discourse-cdn.com/julialang/original/3X/c/e/ce2ed6157364f3a649b1d74551edf1a1211b9e94.png)

---

<div class="post-metadata">

### Author: ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)
#### Post date: [January 31, 2018, 5:46pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/2 "2018-01-31T17:46:53Z")

</div>

Maybe this post can go well with the tutorial that will happen on youtube : [Intro to JuliaDB, a package for working with large persistent data sets - YouTube](https://www.youtube.com/watch?v=d5SzUh2_ono)

which I will definitely watch

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [January 31, 2018, 8:10pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/3 "2018-01-31T20:10:54Z")

</div>

Shameless self-promotion: I’ve also tried to put together a [small WIP tutorial](https://github.com/piever/JuliaDBTutorial/blob/master/hflights.ipynb) on JuliaDB (focused on data manipulation and visualization for researchers - not focused at all on the machine learning side).

It’s based on a R dplyr tutorial (which somebody has already ported to DataFrames [here](https://www.juliabloggers.com/data-wrangling-in-julia-based-on-dplyr-flights-tutorials/). I’ve added something on visualizations which was not present in the original dplyr tutorial but I think it’s relevant topic. Feel free to write feedback in the “Issues” of the repositories if you wish (not to divert this thread from the original topic too much).

@joshday I unfortunately have not included OnlineStats integration for two reasons:

- The recipes PR is not merged yet
- I couldn’t get it to work smoothly with missing data (which I had in my example dataset)  
Still, if you manage to circumvent these issues and want to display some of your really cool work, feel free to modify the notebook and open a PR.

---

<div class="post-metadata">

### Author: ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)
#### Post date: [January 31, 2018, 10:49pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/4 "2018-01-31T22:49:54Z")

</div>

@piever Thanks for the pointer to your tutorial! I’ll respond here partially because dealing with missing data via OnlineStats is a newer feature I’d like to advertise.

1. You can filter and transform data with (as an example):

2. While the plot recipe PR is not in JuliaDB yet, all the functionality is currently available through [`OnlineStats.Partition`](http://joshday.github.io/OnlineStats.jl/latest/visualizations.html#Partitions-1) and [`OnlineStats.IndexedPartition`](http://joshday.github.io/OnlineStats.jl/latest/visualizations.html#Indexed-Partitions-1). What we are adding through the PR is using these two types with a simpler syntax.

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [January 31, 2018, 11:22pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/5 "2018-01-31T23:22:58Z")

</div>

Thanks for the clarification, I’ll definitely try and explore a bit more the JuliaDB-OnlineStats integration

---

<div class="post-metadata">

### Author: ![shashi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shashi/32/1824_2.png) [@shashi](https://discourse.julialang.org/u/shashi)
#### Post date: [February 1, 2018, 6:36am UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/6 "2018-02-01T06:36:59Z")

</div>

We will definitely be covering these topics! Stay tuned.

---

<div class="post-metadata">

### Author: ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)
#### Post date: [February 1, 2018, 6:55am UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/7 "2018-02-01T06:55:07Z")

</div>

A compare and constrast with SAS would be nice

---

<div class="post-metadata">

### Author: ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)
#### Post date: [February 1, 2018, 7:13am UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/8 "2018-02-01T07:13:10Z")

</div>

Looks really cool, thanks for this great work.

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [February 1, 2018, 12:36pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/9 "2018-02-01T12:36:08Z")

</div>

> [@joshday](#):
>
> s = series(Mean(), Variance(); filter = isfinite, transform = abs)  
> reduce(s, table; select = :mycolumn)

I actually couldn’t get it to work with either type of missing data, I think some signatures are maybe too strict (or correct typing is enforced too early in the pipeline) but it’d be definitely really cool to get this to work.

```julia
using OnlineStats
using DataValues
s = series(Mean(), Variance(); filter = !isnull, transform = get)
y = cumsum(randn(10^6)) + 100randn(10^6)
ym = DataValueArray(y)
ym[2] = DataValue()
fit!(s, ym)

```

gives:

```julia

MethodError: no method matching fit!(::OnlineStats.AugmentedSeries{0,OnlineStats.Series{0,Tuple{OnlineStats.Mean,OnlineStats.Variance},OnlineStatsBase.EqualWeight},Base.##57#58{Base.#isnull},Base.#get,Base.#identity}, ::DataValues.DataValue{Float64})
Closest candidates are:
  fit!(::OnlineStats.HyperLogLog, ::Any, ::Float64) at /home/pietro/.julia/v0.6/OnlineStats/src/stats/stats.jl:314
  fit!(::StatsBase.StatisticalModel, ::Any...) at /home/pietro/.julia/v0.6/StatsBase/src/statmodels.jl:104
  fit!(::OnlineStats.CountMap{T}, ::T, ::Float64) where T at /home/pietro/.julia/v0.6/OnlineStats/src/stats/stats.jl:122
  ...

Stacktrace:
 [1] fit!(::OnlineStats.AugmentedSeries{0,OnlineStats.Series{0,Tuple{OnlineStats.Mean,OnlineStats.Variance},OnlineStatsBase.EqualWeight},Base.##57#58{Base.#isnull},Base.#get,Base.#identity}, ::DataValues.DataValueArray{Float64,1}) at /home/pietro/.julia/v0.6/OnlineStats/src/series.jl:177
 [2] include_string(::String, ::String) at ./loading.jl:522

```

and:

```julia
using Missings
s = series(Mean(), Variance(); filter = !ismissing)
y = cumsum(randn(10^6)) + 100randn(10^6)
ym = allowmissing(y)
ym[2] = missing
fit!(s, ym)

```

gives:

```julia
MethodError: no method matching fit!(::OnlineStats.AugmentedSeries{0,OnlineStats.Series{0,Tuple{OnlineStats.Mean,OnlineStats.Variance},OnlineStatsBase.EqualWeight},Base.##57#58{Missings.#ismissing},Base.#identity,Base.#identity}, ::Missings.Missing)
Closest candidates are:
  fit!(::OnlineStats.HyperLogLog, ::Any, ::Float64) at /home/pietro/.julia/v0.6/OnlineStats/src/stats/stats.jl:314
  fit!(::StatsBase.StatisticalModel, ::Any...) at /home/pietro/.julia/v0.6/StatsBase/src/statmodels.jl:104
  fit!(::OnlineStats.CountMap{T}, ::T, ::Float64) where T at /home/pietro/.julia/v0.6/OnlineStats/src/stats/stats.jl:122
  ...

Stacktrace:
 [1] fit!(::OnlineStats.AugmentedSeries{0,OnlineStats.Series{0,Tuple{OnlineStats.Mean,OnlineStats.Variance},OnlineStatsBase.EqualWeight},Base.##57#58{Missings.#ismissing},Base.#identity,Base.#identity}, ::Array{Union{Float64, Missings.Missing},1}) at /home/pietro/.julia/v0.6/OnlineStats/src/series.jl:177
 [2] include_string(::String, ::String) at ./loading.jl:522

```

whereas if everything is `Float64`, it filters just fine:

```julia
s = series(Mean(), Variance(); filter = isfinite)
y = cumsum(randn(10^6)) + 100randn(10^6)
y[2] = NaN
fit!(s, y) # gives correct finite result

```

Worth opening an issue or am I doing something silly?

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [February 1, 2018, 12:50pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/10 "2018-02-01T12:50:22Z")

</div>

There may also be extra typing issues as to filter you need an `AugmentedSeries` whereas the recipe is just for `Series`:

```julia
y = cumsum(randn(10^6)) + 100randn(10^6)
o = Partition(Hist(50))
s = series(y, o, filter = isfinite)
plot(s, xlab = "Nobs")

```

```julia
No user recipe defined for OnlineStats.AugmentedSeries{0,OnlineStats.Series{0,Tuple{OnlineStats.Partition{0,OnlineStats.Hist{OnlineStats.AdaptiveBins{Float64}}}},OnlineStatsBase.EqualWeight},Base.#isfinite,Base.#identity,Base.#identity}

```

---

<div class="post-metadata">

### Author: ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)
#### Post date: [February 1, 2018, 1:31pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/11 "2018-02-01T13:31:43Z")

</div>

All of your examples work for me…but that’s because I forgot to tag an important change in OnlineStatsBase that allows DataValues/Missings as input. There’s a pending PR in Metadata for OnlineStats that let’s you plot AbstractSeries as well.

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [February 1, 2018, 2:11pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/12 "2018-02-01T14:11:41Z")

</div>

Everything works smoothly after `Pkg.checkout`, and syntax will get even better after the transition to `missing` as one no longer needs `transform = get`. Once again, kudos for the amazing work!

---

<div class="post-metadata">

### Author: ![shashi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shashi/32/1824_2.png) [@shashi](https://discourse.julialang.org/u/shashi)
#### Post date: [February 1, 2018, 6:01pm UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/13 "2018-02-01T18:01:41Z")

</div>

I would also like to draw attention to the out-of-core functionality [http://juliadb.org/latest/manual/out-of-core.html](http://juliadb.org/latest/manual/out-of-core.html) it’s limited yet useful for big datasets, especially in combination with OnlineStats.

Cheers!

---

<div class="post-metadata">

### Author: ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)
#### Post date: [August 7, 2018, 9:01am UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/14 "2018-08-07T09:01:26Z")

</div>

Is there a timeline for updating JuliaDB for v0.7? (Not urging anything, I am grateful for free software, just looking for information). I considered making a PR, but there are already multiple ones.

---

<div class="post-metadata">

### Author: ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)
#### Post date: [August 7, 2018, 9:07am UTC](https://discourse.julialang.org/t/ann-new-and-improved-juliadb/8719/15 "2018-08-07T09:07:07Z")

</div>

I know @shashi has been working on it, but I won’t speak for him on the timeline. The first step is IndexedTables ([https://github.com/JuliaComputing/IndexedTables.jl/pull/182](https://github.com/JuliaComputing/IndexedTables.jl/pull/182)), which has a CI-passing PR.
