# Groupedhist with missing values

**URL:** <https://discourse.julialang.org/t/groupedhist-with-missing-values/124792>\
**Category:** New to Julia\
**Tags:** statsplots\
**Created:** [January 15, 2025, 2:04pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792 "2025-01-15T14:04:08Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![mocalvao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mocalvao/32/19318_2.png) [@mocalvao](https://discourse.julialang.org/u/mocalvao)\
**Post date:** [January 15, 2025, 2:04pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/1 "2025-01-15T14:04:08Z")

</div>

Hi,  
I have a dataframe df with several columns, among which :P1, whose eltype is Union{Missing,Float64} and :Nome\_Turma, whose eltype is String. I want to make a groupedhist, from StatsPlots.jl (or otherwise), but I get:

```julia
@df df groupedhist(:P1, group=:Nome_Turma)
ERROR: TypeError: non-boolean (Missing) used in boolean context

```

I know my column df.P1 does have missing values. Might this be the problem? If so, how should I correct it? Shouldn’t the command groupedhist do this automatically?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 15, 2025, 2:44pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/2 "2025-01-15T14:44:49Z")

</div>

Probably

```julia
@df dropmissing(df, :P1) groupedhist(:P1, group = :Nome_Turma)

```

---

<div class="post-metadata">

**Author:** ![mocalvao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mocalvao/32/19318_2.png) [@mocalvao](https://discourse.julialang.org/u/mocalvao)\
**Post date:** [January 15, 2025, 2:49pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/3 "2025-01-15T14:49:40Z")

</div>

> [@nilshg](#):
>
> `@df dropmissing(df, :P1) groupedhist(:P1, group = :Nome_Turma)`

Excellent and on spot!  
Another, perhaps unrelated, question is: how could I generate the histograms in distinct (sub)plots?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 15, 2025, 2:51pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/4 "2025-01-15T14:51:33Z")

</div>

```julia
p1 = histogram(df[df.Nome_Turma .== some_value, :P1])
p2 = histogram(df[df.Nome_Turma .== other_value, :P1])
plot(p1, p2)

```

or something like

```julia
plot([histogram(df[df.Nome_Turma .== x, :P1]) for x in unique(df.Nome_Turma)]...)

```

if there’s lots of values

---

<div class="post-metadata">

**Author:** ![mocalvao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mocalvao/32/19318_2.png) [@mocalvao](https://discourse.julialang.org/u/mocalvao)\
**Post date:** [January 15, 2025, 2:53pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/5 "2025-01-15T14:53:20Z")

</div>

> [@nilshg](#):
>
> `plot([histogram(df[df.Nome_Turma .== x, :P1]) for x in unique(df.Nome_Turma)]...)`

Wow! Thank you so much!!!

---

<div class="post-metadata">

**Author:** ![mocalvao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mocalvao/32/19318_2.png) [@mocalvao](https://discourse.julialang.org/u/mocalvao)\
**Post date:** [January 17, 2025, 1:15pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/6 "2025-01-17T13:15:25Z")

</div>

Perhaps I should create another post, but here it goes: What if I wanted to use the argument `:P1` as a variable inside a loop, such as, e.g.

```julia
etapas = [:P1, :P2, :P3, :SC]
for etapa in etapas
  ghist = @df dropmissing(df, etapa) groupedhist(etapa, group = :Nome_Turma)
  savefig(ghist, "ghist_$etapa")
end

```

It issues an error:

```julia
ERROR: MethodError: no method matching groupedvec2mat(::Dict{Int64, Int64}, ::Vector{Int64}, ::String, ::RecipesPipeline.GroupBy, ::Float64)

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 17, 2025, 2:12pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/7 "2025-01-17T14:12:36Z")

</div>

Yeah that won’t work because `@df` is a macro which can only see the actual code written, not any runtime values. This has nothing to do with the `dropmissing`, you just can’t write `@df df groupedhist(etapa, group = :Nome_Turma)` as the macro will expand this using the actual string `etapa` rather than the value of that variable.

---

<div class="post-metadata">

**Author:** ![mocalvao](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mocalvao/32/19318_2.png) [@mocalvao](https://discourse.julialang.org/u/mocalvao)\
**Post date:** [January 17, 2025, 2:40pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/8 "2025-01-17T14:40:01Z")

</div>

> [@nilshg](#):
>
> plot([histogram(df[df.Nome\_Turma .== x, :P1]) for x in unique(df.Nome\_Turma)]…)

From the JuliaPlots/StatsPlots.jl site I thought

```julia
using StatsPlots
etapa = :P3
@df dropmissing(df, cols(etapa)) groupedhist(cols(etapa), group = :Nome_Turma)

```

should work but the following error was issued:

```julia
ERROR: UndefVarError: `cols` not defined in `Main`

```

Any suggestions?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 17, 2025, 3:13pm UTC](https://discourse.julialang.org/t/groupedhist-with-missing-values/124792/9 "2025-01-17T15:13:57Z")

</div>

That doesn’t change the fact that you want the macro to understand the _runtime value_ of a variable, which just isn’t possible. Macros are just convenience to rewrite code, and the `@df` macro needs an actual Symbol in the written expression it transforms to work. See these two examples (where I’ve replaced your column names `P1` with `b` and `Nome_Turma` with `group`):

```julia
julia> prettify(@macroexpand(@df df groupedhist(:b, group = :group)))
:(((fly->begin
          ((penguin, locust), grasshopper) = (StatsPlots).extract_columns_and_names(fly, :b, :group)
          (StatsPlots).add_label(["b"], groupedhist, penguin, group = locust)
      end))(df))

julia> prettify(@macroexpand(@df df groupedhist(etapa, group = :group)))
:(((fly->begin
          ((penguin,), locust) = (StatsPlots).extract_columns_and_names(fly, :group)
          (StatsPlots).add_label(["etapa"], groupedhist, etapa, group = penguin)
      end))(df))

```
