# Dealing with different concepts of "missingness"

**URL:** <https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008>\
**Category:** New to Julia\
**Created:** [February 10, 2021, 4:31pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008 "2021-02-10T16:31:16Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 4:31pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/1 "2021-02-10T16:31:16Z")

</div>

I’ve come accross different types of missingness that occur when loading and preprocessing data.

1. `NaN`
2. `Nothing`
3. `Missing`

e.g.

```julia
df = groupby(DataFrame(g = [1,1,2], v = [1,2,3]), :g) |>
          x-> combine(x, :v =>var)
df.v_var[2] == NaN
json = JSON3.read("""{"Nothing":null}""")
isnothing(json["Nothing"])
1 + missing == missing

```

This brings some problems with processing the data e.g. Arrow.jl does not play nice with the `Nothing` type, coalesce does not work with `NaN`. I wonder how you are dealing with these differnt types?

Thanks

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 10, 2021, 4:41pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/2 "2021-02-10T16:41:43Z")

</div>

1. Don’t use `nothing` with data. `nothing` is for functions that don’t return anything, or for undefined returns, i.e. failed regex matches.
2. Use `missing` with data. `missing` means, “here is a value, we just don’t know what it is”
3. `NaN` is a `Float64`, it’s a number. It’s not quite right to say you “don’t know what it is”, because it’s really the result of some kind of “bad” numerical calculation, i.e. `0/0`.

`coalesce` is just for `missing`s. The function `something` is for `nothing` values. Looking at it now this is not the best naming scheme, I guess.

I don’t think there exists an equivalent function for `NaN`.

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 4:59pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/3 "2021-02-10T16:59:06Z")

</div>

Thanks for the clarification! I’ve been using `replace` quite a lot since it seems to work with all of them. But it’s kind of annoying to deal with all of them depending on the package/function used. I already wrote a [small function](https://github.com/danielw2904/DFUtils.jl/blob/master/src/DFUtils.jl#L153) that deals with `nothing` but I guess I have to add `NaN` .

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 10, 2021, 5:03pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/4 "2021-02-10T17:03:06Z")

</div>

Just so you know, with that function you wrote, you aren’t actually avoiding any copying.

You could just do

```julia
df[!, c] = replace(df[!, c], nothing => missing)

```

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 5:12pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/5 "2021-02-10T17:12:33Z")

</div>

Good to know! I thought that this was an in-place operation. Is there a way to do it without copying or is your suggestion the way?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 10, 2021, 5:16pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/6 "2021-02-10T17:16:07Z")

</div>

No, there isn’t a to change the type of a vector to allow `missing`s without copying. Maybe in the future that could be an optimization, but your `Convert` call is where the copying happens, I think.

---

<div class="post-metadata">

**Author:** ![lostella](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lostella/32/356_2.png) [@lostella](https://discourse.julialang.org/u/lostella)\
**Post date:** [February 10, 2021, 5:31pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/7 "2021-02-10T17:31:27Z")

</div>

> [@pdeffebach](#):
>
> Don’t use `nothing` with data. `nothing` is for functions that don’t return anything, or for undefined returns, i.e. failed regex matches.

There could be cases where fields are optional, and `nothing` should be used, right? “I know that there is _no_ value here”

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 10, 2021, 5:39pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/8 "2021-02-10T17:39:43Z")

</div>

Yes. But in the context of data I’m not sure that’s a very common scenario to be in. You never see `NULL` with working with data frames in R for example.

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 6:42pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/9 "2021-02-10T18:42:23Z")

</div>

I would agree. Coming from R I have never seen NULL used in tabular data. But maybe my usecase is specific in that I think in terms of a model matrix s.t. it does not really matter _why_ a value is missing. In any case the row will be removed or I would like to replace it with some value.

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 6:54pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/10 "2021-02-10T18:54:35Z")

</div>

The problem is that one still has to check whether there are `nothing`s in the column otherwise missings are allowed implicitly

```julia
julia> typeof(replace([1,2,3], nothing => missing))
Array{Union{Missing, Int64},1}

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 10, 2021, 6:57pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/11 "2021-02-10T18:57:11Z")

</div>

Yeah, you still need that check.

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [February 10, 2021, 7:00pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/12 "2021-02-10T19:00:13Z")

</div>

My lazy version of that is

```julia
if eltype(df[!, col]) isa Union     
...
end

```

The more involved version to read the types of a union into a vector

```julia
function readtypes(U::Union, types = DataType[])
    push!(types, getfield(U, :a))
    if isa(getfield(U, :b), DataType)
        push!(types, getfield(U, :b))
        return types
    else
        readtypes(U.b, types)
    end
end

```

---

<div class="post-metadata">

**Author:** ![Nathan\_Boyer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nathan_boyer/32/14825_2.png) [@Nathan\_Boyer](https://discourse.julialang.org/u/Nathan_Boyer)\
**Post date:** [February 10, 2021, 7:07pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/13 "2021-02-10T19:07:06Z")

</div>

I asked a [similar question on Stack Overflow](https://stackoverflow.com/questions/61936371/usage-and-convention-differences-between-missing-nothing-undef-and-nan-in-jul) that received some helpful answers you may find useful.

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [February 10, 2021, 7:10pm UTC](https://discourse.julialang.org/t/dealing-with-different-concepts-of-missingness/55008/14 "2021-02-10T19:10:20Z")

</div>

If your goal is just strip away any `missing`, `nothing`, and `NaN` entries …

```julia
notmissing = Base.Fix2(!==, missing)
notnothing = Base.Fix2(!==, nothing)
notnan(x) = true
notnan(x::Base.IEEEFloat) = !isnan(x)
isavailable(x) = notmissing(x) && notnothing(x) && notnan(x)
clean(data) = filter(isavailable, data)

```

If you want the indices of each `missing`, `nothing`, and `NaN`

```julia
function unavailable(data)
  idxs = []
  for (i,x) in enumerate(data)
    !isavailable(x) && push!(idxs, i)
  end
  return idxs
end

```
