# What workflows for missing values are more ergonomic in Julia?

**URL:** <https://discourse.julialang.org/t/what-workflows-for-missing-values-are-more-ergonomic-in-julia/106923>\
**Category:** Internals & Design\
**Created:** [November 30, 2023, 3:00am UTC](https://discourse.julialang.org/t/what-workflows-for-missing-values-are-more-ergonomic-in-julia/106923 "2023-11-30T03:00:37Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [November 30, 2023, 3:00am UTC](https://discourse.julialang.org/t/what-workflows-for-missing-values-are-more-ergonomic-in-julia/106923/1 "2023-11-30T03:00:37Z")

</div>

Continuing the discussion from [Why are missing values not ignored by default?](https://discourse.julialang.org/t/why-are-missing-values-not-ignored-by-default/106756/225):

Hi, all, we started talking about some issues about dropping missing automatically, in the discussion it became clear that some people tend not to experience the same level of non-ergonomic issues as others.

What workflow recommendations can we come up with that we could offer to new users to make Julia’s existing missing handling be as helpful as possible? What kinds of things should they avoid? Imagine we are writing some course notes for a data analysis undergrad class or something…

---

<div class="post-metadata">

**Author:** ![Tetrakai](https://avatars.discourse-cdn.com/v4/letter/t/4da419/32.png) [@Tetrakai](https://discourse.julialang.org/u/Tetrakai)\
**Post date:** [November 30, 2023, 4:27pm UTC](https://discourse.julialang.org/t/what-workflows-for-missing-values-are-more-ergonomic-in-julia/106923/2 "2023-11-30T16:27:15Z")

</div>

Writing skipmissing everywhere sounds annoying. But I think a package should do it, and _not_ overwrite the default function. It could be mean\*(vec) or whatever.

But when I deal with missings, usually they contain information. So have mean\*() return the mean optionally along with the number/proportion of missings and invalids.

Ie:

`mean([11, 2, 1, 10, missing])` is not the same to me as `mean([missing, 2, missing, 10, missing])`, which is different from `mean([missing, 2, missing, 10, "6"])`

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [November 30, 2023, 5:59pm UTC](https://discourse.julialang.org/t/what-workflows-for-missing-values-are-more-ergonomic-in-julia/106923/3 "2023-11-30T17:59:48Z")

</div>

> [@Tetrakai](#):
>
> _not_ overwrite the default function. It could be mean\*(vec) or whatever.

I just tend to make a shorter alias like `sm` and write `mean(sm(vec))`, maybe alias a composition of a function with `sm` if it happens enough. I could never make an “automatic” `skipmissing` work because I don’t always want to `skipmissing`, so I might as well make the case-by-case basis shorter to write. If I have to exclude rows with `missing`s across multiple select columns, lazy `dropmissing` on the overall dataframe it is.

In that thread, I used 1 higher order function to refactor a set of variant scalar operations with 1 particular way of replacing the original operation with `false` and propagating `missing` for the other. I should note that as scalar operations, they do not do anything like skipping missings, I’m just mentioning that higher order functions can be used here too.
