# Plotting options for DataFrames with NullableArray

**URL:** <https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584>\
**Category:** Visualization\
**Tags:** question\
**Created:** [January 19, 2017, 10:18am UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584 "2017-01-19T10:18:53Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [January 19, 2017, 10:18am UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/1 "2017-01-19T10:18:53Z")

</div>

I have been transitioning to `Nullable`s in my data analysis code, then belatedly realized that neither Gadfly nor Plots.jl work with current master DataFrames. Is there a plotting package that is usable at the moment?  
[https://github.com/GiovineItalia/Gadfly.jl/issues/909](https://github.com/GiovineItalia/Gadfly.jl/issues/909)  
[https://github.com/JuliaPlots/StatPlots.jl/issues/33](https://github.com/JuliaPlots/StatPlots.jl/issues/33)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [January 19, 2017, 10:48am UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/2 "2017-01-19T10:48:13Z")

</div>

Here is a workaround:

```julia
"""
Remove missing values from a dataframe, and make sure that all the
columns are plain vanilla Vectors, not NullableVectors.

Horrible hack, until plotting libraries catch up.
"""
function sanitize_df(df)
    df = df[complete_cases(df),:]
    colnames = names(df)
    _column(name) = [get(x) for x in df[name]]
    ## relying on https://github.com/JuliaStats/DataFrames.jl/issues/1119
    ## in the spirit of https://xkcd.com/1172/
    DataFrame(map(_column, colnames), colnames)
end

```

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 19, 2017, 4:09pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/3 "2017-01-19T16:09:27Z")

</div>

Why do you think that is a horrible hack? StatPlots have always assumed DataFrames had complete cases (the alternative might be to convert NAs/Nulls to NaN when the eltype can be promoted to Float64), and in fact your suggestion is pretty much what is recommended by @nalimilan here: [https://github.com/JuliaStats/DataFrames.jl/issues/1148](https://github.com/JuliaStats/DataFrames.jl/issues/1148)  
EDIT: I see now. It is ‘horrible’ because you rely on this DataFrame constructor that avoids Nullables [https://github.com/JuliaStats/DataFrames.jl/blob/bf0bda80a1c4e24f0fdc0a547883ea6e285de53d/src/dataframe/dataframe.jl#L51](https://github.com/JuliaStats/DataFrames.jl/blob/bf0bda80a1c4e24f0fdc0a547883ea6e285de53d/src/dataframe/dataframe.jl#L51)

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [January 19, 2017, 4:44pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/4 "2017-01-19T16:44:14Z")

</div>

I guess another issue with this approach is that some more or less irrelevant column, which is not used in the plot, could have plenty of missing values. You don’t want to lose data because of that but also you don’t want to think about which columns are involved every time you do a plot.  
At least as far as StatPlots is concerned, it seems like one could do better. [This macro](https://github.com/JuliaPlots/StatPlots.jl/blob/master/src/dataframes.jl#L39) replacing a dataframe and a list of symbols with the respective columns could also discard all lines where at least one of the _involved_ columns has a NULL and replace each symbol with the respective column converted to array.

The whole `groupapply` stuff is probably a bit trickier, but I’m waiting for this DataFrames situation to stabilize to understand what needs to be done. Also there, a simple solution would be to add a few lines of code at the beginning which discard lines where in one relevant column there’s a NULL. My only worry is that, even when starting with a dataframe of arrays, maybe some operations (say `by` or subselecting the DataFrame, or adding columns) could restart producing NullableArrays. To be honest though, I really hope that there will be a way to simply avoid dealing with NullableArrays when you don’t have missing data.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 19, 2017, 4:50pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/5 "2017-01-19T16:50:00Z")

</div>

Yes, exactly, my suggestion ATM is to use that recipe to discard all non-complete cases of those variables that are referenced inside the call and then replace the variables with normal arrays. The group\_apply recipe should be able to use the same function to handle the df.

---

<div class="post-metadata">

**Author:** ![tbreloff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbreloff/32/68_2.png) [@tbreloff](https://discourse.julialang.org/u/tbreloff)\
**Post date:** [January 19, 2017, 5:06pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/6 "2017-01-19T17:06:30Z")

</div>

I have a couple of things to say:

- For StatPlots, the proper solution is likely to have a “type recipe” for NullableArray, which does a consistent conversion of some sort to a type that is plottable. Then a DataFrame can be decomposed into a bunch of NullableArrays, and the type recipe can handle further conversions
- Removing rows that contain nulls is likely not ideal. There are plenty of times when missing data can (and should) have an affect on the visualization. The “correct” way to deal with this is likely to map to NaN for numeric data, and to add native handling of Nullables when processing discrete inputs within Plots (which internally would add NaN to the numeric mapping discrete–\>continuous)

Of course, getting the no-null case working should be the first priority, so any hack is fine.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 19, 2017, 6:29pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/7 "2017-01-19T18:29:12Z")

</div>

- Why not just update the current type recipe for `DataFrames`, replacing `collect(df[:x])` calls with `get` calls (and the other handling described here (no need to use the ‘horrible’ hack of rebuilding the DataFrame))?
- Yes, exactly, that is what I meant by converting to NaN for types that can be promoted to Float64. I probably didn’t make it clear enough, nor realized that discrete types could be replaced by NaN inside Plots.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 19, 2017, 10:20pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/8 "2017-01-19T22:20:14Z")

</div>

@Tamas_Papp please check out and comment on the Nullable\_DataFrames branch of StatPlots 🙂 It should work for all cases except for when the DataFrame column passed to `plot` is not numeric AND contains Nulls

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [January 20, 2017, 11:52am UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/9 "2017-01-20T11:52:37Z")

</div>

EDIT: missed the latest commit. Now it works fine.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 20, 2017, 8:08pm UTC](https://discourse.julialang.org/t/plotting-options-for-dataframes-with-nullablearray/1584/10 "2017-01-20T20:08:26Z")

</div>

Cool, that is the current recommendation for plotting DataFrames with Nullables with Plots, then.
