# Why are missing values not ignored by default?

**URL:** <https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756>\
**Category:** Internals & Design\
**Tags:** data, missing-values\
**Created:** [November 26, 2023, 1:45am UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756 "2023-11-26T01:45:36Z")\
**Posts on this page:** 20\
**Page:** 3

<div class="post-metadata">

**Author:** ![adienes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adienes/32/37459_2.png) [@adienes](https://discourse.julialang.org/u/adienes)\
**Post date:** [November 27, 2023, 2:47pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/43 "2023-11-27T14:47:48Z")

</div>

> [@sijo](#):
>
> What do you mean it’s false?

because `missing`, as a value, is not equal to any other value in any meaningful sense

I understand the idea behind three-valued logic, but I wish this had been opt-in in something like `ThreeValuedLogic.jl` instead of baked into Base into functions as basic as `==`

> Imagine all the computations that would … give … results

I imagine this frequently 🙃

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [November 27, 2023, 3:29pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/44 "2023-11-27T15:29:37Z")

</div>

Would you agree that `x < y` implies `y > x`?

If `missing < y` is `false` , then `y < missing` should be true, but that contradicts even your expectations, no?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 3:32pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/45 "2023-11-27T15:32:20Z")

</div>

You can have both `missing < y` be `false` and `missing > y` be `false`. It’s not a number, it’s a `missing` value and doesn’t need to have the same semantics as the reals.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [November 27, 2023, 3:47pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/46 "2023-11-27T15:47:44Z")

</div>

Seems to me from the discussion that the best would be to have a package that deals with another type, such as `unsafe_missing`, that implemented all the functions desired, so that someone can just replace all the `missing`s in their data with that one and move forward with their preferred choice.

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [November 27, 2023, 3:56pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/47 "2023-11-27T15:56:59Z")

</div>

I see. But if `missing` doesn’t need to have the same semantics as real numbers, why mix that semantics with the semantics of the reals?

Let’s say we skip `missing`s by default (which the original discussion is about). Then we have `x + missing = x` (e.g., `missing` acts as “zero” here). What about `x * missing`? Does `missing` still act as zero here? That would not align with the “skip missing” expectations, right? So maybe it should act as one, i.e., `x * missing = x`. But then we loose, e.g., the distributive property `a * (b + c) = a * b + a * c` because `a * (b + missing) = a * b` while `a * b + a * missing = a * b + a`.

So yeah, imho the current behavior is only reasonable. The user should decide the semantics of `missing`.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [November 27, 2023, 3:58pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/48 "2023-11-27T15:58:54Z")

</div>

> [@lmiq](#):
>
> Seems to me from the discussion that the best would be to have a package that deals with another type, such as `unsafe_missing`, that implemented all the functions desired, so that someone can just replace all the `missing`s in their data with that one and move forward with their preferred choice.

It would make it easier to do something like that if we added a new supertype `AbstractMissing >: Missing` to the `Base` package. (This can be done in 1.x — adding a new supertype is not considered a breaking change IIRC.)

That way, a lot of packages like DataFrames.jl could more easily support alternative `missing` types, e.g. an `ImputedMissing` type with some kind of semantics of automatic imputation, without having to specifically depend on new packages.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 4:52pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/49 "2023-11-27T16:52:13Z")

</div>

1. The only Pareto improvement at the moment is to figure out better nicknames for common operations. `skipmissing` is simply too long for people to want to type. I think everyone would be happier if we had

2. Making a new `missing` value is not a good avenue. If we make `IgnoreMissing` a type in a new package, we would have to have overloads for _a lot_ of statistical functions in Base or Statistics, like `mean`, `cor`, `sum`, etc. It’s one thing to add `missing` support for base functions, but it adding them for _new_ packages would involve convincing every author of a new statistical package to support the `IgnoreMissing` type. Given the responses in this thread, many authors do not believe `missing`s should ever be skipped. It would involve having the same (tired) arguments about `missing` semantics over and over again for every package. In practice, `IgnoreMissing` would simply error or require it’s own `skipmissing` function, making it indistinguishable from the existing problems.

3. Many of the solutions proposed in this thread are not realistic solutions to the challenges people working with messy data face.

4. I am very sympathetic to Julia developers and those who write packages and how `missing` propagation can make it more difficult to maintain large code-bases. Someone doing data analysis might prefer `missing < y` to be `false` (the semantics being, `<` is only true if we are _sure_ the value is less than `y`, which is perfectly acceptable), but someone maintaining, say, the VS Code extension will not want that behavior.

Given these trade-offs, I want to re-iterate point 1. We simply need better syntax for handling `missing` values. The names `passmissing` and `skipmissing` should be shortened, and we should write more specialized methods for common functions to ensure summations are correct.

[DataFramesMeta.jl](https://github.com/JuliaData/DataFramesMeta.jl) provides some features to make this easier, such as dropping missing values in `@subset`, adding the `@passmissing` macro-flag, and hopefully more in the future.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 27, 2023, 5:04pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/50 "2023-11-27T17:04:34Z")

</div>

> [@stevengj](#):
>
> It would make it easier to do something like that if we added a new supertype `AbstractMissing >: Missing` to the `Base` package. (This can be done in 1.x — adding a new supertype is not considered a breaking change IIRC.)
> 
> That way, a lot of packages like [DataFrames.jl](https://juliahub.com/ui/Packages/DataFrames) could more easily support alternative `missing` types, e.g. an `ImputedMissing` type with some kind of semantics of automatic imputation, without having to specifically depend on new packages.

I’ve experimented custom missings values (but with a different goal) in [TypedMissings.jl](https://github.com/nalimilan/TypedMissings.jl), but it’s currently blocked by some hard issues in Julia:

> <https://github.com/JuliaLang/julia/pull/44407>
>
> This is more generic and will allow packages to define custom \`missing\`-like typ…es allowing to distinguish several kinds of
> missing values like in e.g. Stata and SAS (see https://github.com/nalimilan/TypedMissings.jl). This should have no performance impact now thanks to #38905.
> 
> https://github.com/JuliaCI/BaseBenchmarks.jl/pull/294 adds more benchmarks to cover changes made in this PR, let's see whether Nanosoldier confirms that no performance regressions happen.

> <https://github.com/JuliaLang/julia/issues/38241>
>
> Consider this code:
> \`\`\`julia
> julia\> bar(::Nothing) = missing
> bar (generic fun…ction with 1 method)
> 
> julia\> bar(x) = x;
> 
> julia\> xs = \[1, 2, nothing, 4\]
> 4-element Vector{Union{Nothing, Int64}}:
> 1
> 2
> nothing
> 4
> 
> julia\> ys = map(bar, xs)
> 4-element Vector{Union{Missing, Int64}}:
> 1
> 2
> missing
> 4
> \`\`\`
> We can see we get a nice array of small unions in the output.
> 
> 
> But if I make my own similar singleton type, then the output is a \`Array{Any}\`.
> 
> \`\`\`julia
> julia\> struct Null end
> 
> julia\> foo(::Nothing) = Null();
> 
> julia\> foo(x) = x;
> 
> julia\> xs = \[1, 2, nothing, 4\]
> 4-element Vector{Union{Nothing, Int64}}:
> 1
> 2
> nothing
> 4
> 
> julia\> ys = map(foo, xs)
> 4-element Vector{Any}:
> 1
> 2
> Null()
> 4
> \`\`\`
> 
> Is there something extra i need to do, or is the compiler cheating for \`Missing\` and \`Nothing\`?

---

<div class="post-metadata">

**Author:** ![Eben60](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eben60/32/13475_2.png) [@Eben60](https://discourse.julialang.org/u/Eben60)\
**Post date:** [November 27, 2023, 5:08pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/51 "2023-11-27T17:08:02Z")

</div>

> [@pdeffebach](#):
>
> I think everyone would be happier if we had

From the blog article by [Kamiński](https://bkamins.github.io/julialang/2022/06/17/missing.html), which I’ve already cited here

> Such opinions are always debatable, so recently I decided to run a small [pool on Julia Discourse](https://discourse.julialang.org/t/should-we-add-sm-as-a-short-form-of-skipmissing/82610) about the `skipmissing` function. The question was if we want to shorten the `skipmissing` name into something that is more convenient to use in interactive work. To my surprise, a vast majority of voters preferred a verbose and explicit operation name.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 5:08pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/52 "2023-11-27T17:08:47Z")

</div>

`sm` is a pretty small name compared to `mskip`. So I would take that poll with a grain of salt.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [November 27, 2023, 5:29pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/53 "2023-11-27T17:29:34Z")

</div>

> [@pdeffebach](#):
>
> On the one hand, the vast majority of users probably want `missing` to `y < missing` to be `false`

This is horrifying to me, `true`/`false` is not missing data to me and I want to distinguish it from missing data.

> [@pdeffebach](#):
>
> You can have both `missing < y` be `false` and `missing > y` be `false`

That would imply the missing data is `==y` in which case it’s not really missing; this would be dangerous if I decided to do `.!((data_vector .< y_cutoff) .|| (data_vector .> y_cutoff))`, probably in separate steps, instead of `data_vector .== y_cutoff`. barucden was implying these silent implications are dangerous and likely why operations on `missing` return `missing`.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 5:31pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/54 "2023-11-27T17:31:47Z")

</div>

I don’t want to get into a discussion about the exact semantics of `missing` in this thread. Obviously reaching agreement on these semantics is difficult, if not impossible. I only emphasize that `missing` is not a number and comparisons don’t need to imply completeness or transitivity.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [November 27, 2023, 5:36pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/55 "2023-11-27T17:36:12Z")

</div>

I think the whole point here was that these decisions (`missing` three-valued logic, whether to skip or not, etc) were intended to be made by the developer, which is consistent with the approach taken in a large number of other cases. So, totally fine if you want to ignore missing values by default, that is your decision to make. Your complaint then is about how easy (or not) that is to accomplish.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [November 27, 2023, 5:45pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/56 "2023-11-27T17:45:08Z")

</div>

> [@pdeffebach](#):
>
> I only emphasize that `missing` is not a number and comparisons don’t need to imply completeness or transitivity.

This and your previous comments are indeed discussing the semantics of `missing`, so I’m not sure how you’d prefer avoiding the topic. If you mean you want to avoid the theory, sure, let’s talk practicality. I definitely think it’s impractical and buggy if operations on `missing` values return anything but `missing` by default. Having a `missing` value is important for skipping them, too. Say I have `[false, true, false]` and want to divide the number of `true`s by the number of comparisons of real data. How could I possibly know which `false` wasn’t computed from missing values?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 5:52pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/57 "2023-11-27T17:52:25Z")

</div>

It’s the data analyst’s job to ensure data integrity. The question “which observations contribute to this statistic” is something Julia, the language, can’t answer. The analyst should absolutely conduct additional robustness checks about how missing values are handled and what’s the appropriate way to deal with them.

The question is whether imposing `skipmissing(...)` or propagation on Boolean operations is the the right way to go about that. It’s costly for users to write `skipmissing` every time they wish to calculate the `mean`. I’m simply making an argument that the cost isn’t always worth the benefits.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [November 27, 2023, 6:05pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/58 "2023-11-27T18:05:06Z")

</div>

Couldn’t someone just write a package called `SkipMissingByDefault` thst implements the following?

```julia
module SkipMissingByDefault
mean(x::AbstractArray; kwargs...) = Statistics.mean(skipmissing(x); kwargs...)
end

```

Then the user can do the following.

```julia
using SkipMissingByDefault as sm
sm.mean(...)

```

Or they could just import it to use the modified `SkipMissingByDefault.mean` rather than `Statistics.mean`.

```julia
using SkipMissingByDefault: mean
mean(...)

```

I’m pretty sure `Statistics.mean` will not change its default behavior. That would be both breaking and likey lead to correctness mistakes. The most I could see would be a new keyword argument for convenience.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [November 27, 2023, 6:08pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/59 "2023-11-27T18:08:19Z")

</div>

> [@pdeffebach](#):
>
> It’s costly for users to write `skipmissing` every time they wish to calculate the `mean`. I’m simply making an argument that the cost isn’t always worth the benefits.

That argument is not parallel to whether `missing`s should propagate something other than `missing`, in fact it undermines it. Say I want the average height in a polo team but only two people showed up `[5.2, missing, 5.6, missing]`. Well I had a yardstick, but everyone else is more comfortable with metric, so I convert feet to centimeters…and the `missing`s get replaced with dummy values `[158.496, 0.0, 170.688, 0.0]` (it could return `false`, but it just gets auto-converted to `Float64`). `SkipMissingByDefault.mean` can’t do any different than `Statistics.mean` anymore. But wait, maybe we can skip all `zero(T)`, after all nobody is 0cm tall. But what about cases where zeros are valid data, or what if a type doesn’t _have_ a zero? It’d be nice if we had a dedicated singleton type to skip in folding operations, whether manually or automatically, and they didn’t vanish in the preceding elementwise scalar operations.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 6:26pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/60 "2023-11-27T18:26:53Z")

</div>

I would consider that approach worse than what we currently have in Missings.jl

`SkipMissingByDefault` would have to write new versions of all Base and Statistics functions, and it wouldn’t be able to handle new third-party packages.

By contrast, `skipmissing` and the future `spreadmissing`s features will allow users to interact with _any_ function that allows iterators or `AbstractVector`s, even those not written yet.

But yes, we do need to write safer methods for `::SkipMissing` in order to protect against floating point summation differences.

Our current approach is very flexible and leverages Julia’s multiple dispatch and functions-as-objects approach nicely.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 27, 2023, 6:31pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/61 "2023-11-27T18:31:54Z")

</div>

I’m not arguing that `missing / x` should evaluate to `false`, which you seem to be implying by this post. I’m discussing a scenario where `x < y` should imply “We are absolutely sure that `x` is less than `y`”. This obviously has some costs, I admit. Also, see my above post where I advocate for macro-based approaches to make this easier.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [November 27, 2023, 6:39pm UTC](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/62 "2023-11-27T18:39:20Z")

</div>

> [@Mason](#):
>
> It’s not a bug, it’s the intended behaviour of `mean` hitting a generic fallback. There are _so_ many different wrapper types out there in the world that specializing all of them to account for `skipmissing` is an unending game of whack-a-mole.

Mostly I think this is just a more generic problem, of poor support for . The problem is Julia uses the iterator interface for everything vector-related (functional primitives like `map`, `filter`, `reduce`, etc. all live in `Iterators.X`), even though most data manipulation in Julia works on vectors, rather than lazy one-at-a-time iterators.

Transducers.jl and OnlineStats.jl fix this, but most people don’t use them because they’re not in the stdlib/StatsBase, respectively.

[Previous page](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756.md?page=2)

[Next page](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756.md?page=4)
