# Is there any reason to use NaN instead of missing?

**URL:** <https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396>\
**Category:** General Usage\
**Tags:** missing-values\
**Created:** [July 18, 2022, 9:58am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396 "2022-07-18T09:58:14Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [July 18, 2022, 9:58am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/1 "2022-07-18T09:58:14Z")

</div>

I know that NaN isn’t going away because of the IEEE 754 standard on floating point behavior, but is there ever a reason to use NaN instead of converting and using `missing`s as much as possible? Both are propagating non-values, only NaN seems to be less good at it e.g. `NaN^0` evaluates to `1.0`. The extra bit of type instability of `Missing` unions seems well-contained by `Missing` propagation and optimized by small Union optimizations, so checking `ismissing` never seemed different from checking `isnan`.

I’ve also heard about two kinds of NaNs in the IEEE standard, quiet vs signalling, but I haven’t seen this distinction in Julia yet. Is that important for anything that `missing` doesn’t cover?

---

<div class="post-metadata">

**Author:** ![SteffenPL](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steffenpl/32/206270_2.png) [@SteffenPL](https://discourse.julialang.org/u/SteffenPL)\
**Post date:** [July 18, 2022, 10:18am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/3 "2022-07-18T10:18:37Z")

</div>

I think the type instability is relevant in many settings, for example

```julia
using BenchmarkTools

x = rand(1000)
y = rand(1000)

y[500] = NaN
z = vcat(y[1:499], missing, y[501:end])

function test_add(x,y)
    for i in 1:1000 
        y[i] = x[i] + y[i] 
    end
    y
end

@btime test_add($x, $y) # 723.846 ns (0 allocations: 0 bytes)
@btime test_add($x, $z) # 966.667 ns (0 allocations: 0 bytes)

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 18, 2022, 11:59am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/4 "2022-07-18T11:59:32Z")

</div>

Yes, there are lots of performance (type stability) reasons to use NaN!  
The overhead with `missing` can be more than an order of magnitude even in this simple example:

```julia
julia> x_nan = [([1., 2., 3.],), ([1., 2., 3., NaN],)]

julia> x_missing = [([1., 2., 3.],), ([1., 2., 3., missing],)]

julia> @btime map(x -> sum.(x), $x_nan)
  35.578 ns (1 allocation: 80 bytes)

julia> @btime map(x -> sum.(x), $x_missing)
  378.745 ns (8 allocations: 240 bytes)

```

---

<div class="post-metadata">

**Author:** ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)\
**Post date:** [July 18, 2022, 2:41pm UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/5 "2022-07-18T14:41:20Z")

</div>

`NaN` and `missing` denote semantically different things. Idiomatically, `NaN` is an indication of an arithmetic failure (eg `0.0*Inf`) or is sometimes used to poison computations. The math rules for `NaN` mostly makes sense if one interprets `NaN` as “this _is_ number, but I don’t know which number it should be” (hence `NaN^0.0==1.0`, since _any_ number to the zero power is one - although `0^0` is sometimes debated). `missing` refers to a value that never existed in the first place or was derived from such a number.

Sometimes, one doesn’t need this semantic distinction, in which case `NaN` can often serve the same role as `missing` (mostly as a poison value).

There will rarely be a performance benefit to using `missing`, but it can be make the code much more sensible in places. That said, the performance penalty is usually not huge (although can be significant in micro-benchmarks). I wouldn’t avoid `missing` for performance reasons if you’re actually trying to represent missing values.

As for quiet `qNaN` versus signaling `sNaN`, Julia does not make any distinction between the two. As far as I understand, Julia will never produce a `sNaN` in typical operation. Calling `reinterpret` is the primary way I’d expect them to arise in Julia. That said, most arithmetic operations on a `sNaN` will produce a `qNaN`, so they don’t tend to propagate very far.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [July 18, 2022, 11:17pm UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/6 "2022-07-18T23:17:19Z")

</div>

Just to summarize this thread.  
NaN represents a real number that is unknown from the input expression. NaN does not intend to represent missing data that always propagates in operations, or even all the indeterminate forms arising from limits:

```julia
julia> NaN^0, 1^NaN, NaN > 0 # nonpropagating cases
(1.0, 1.0, false)

julia> 0/0, Inf/Inf, 0*Inf, Inf-Inf, 0^0, 1^Inf, Inf^0 # indeterminate forms
(NaN, NaN, NaN, NaN, 1, 1.0, 1.0)

```

I should use `missing` and `skipmissing` when it’s proper and accept some performance hits as necessary. And if it helps performance to condense non-missing data, I could use `collect` or `append!` on `Base.SkipMissing` instances.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [July 19, 2022, 12:22am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/7 "2022-07-19T00:22:59Z")

</div>

NaNs are very often used to represent missing data in grids. For grids, it makes no sense to have them replace with `missing` specially if one want to save those grids in disk files where libraries like netCDF have no idea what a `missing` is.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [July 19, 2022, 1:04am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/8 "2022-07-19T01:04:17Z")

</div>

> [@mikmoore](#):
>
> hence `NaN^0.0==1.0`, since _any_ number to the zero power is one

IMO, this makes no sense. **NaN** stands for **N** ot **aN** umber so `NaN^0` should be NaN, not 1

In Matlab

```julia
>> NaN^0

ans =

   NaN

```

---

<div class="post-metadata">

**Author:** ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)\
**Post date:** [July 19, 2022, 1:22am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/9 "2022-07-19T01:22:54Z")

</div>

According to [what wikipedia says about IEEE754-2008](https://en.wikipedia.org/wiki/NaN#Function_definition), there are several different versions of the power function that should have different `NaN` behavior. One of which is that `qNaN^0.0` should return `1.0`.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [July 19, 2022, 1:39am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/10 "2022-07-19T01:39:41Z")

</div>

Quoting section 9.2.1 of IEEE 754-2008: “pow(x, ±0) is 1 for any x (even a zero, quiet NaN, or infinity)”. Ditto for “pown(x, 0)”, which is for integer exponents. Maybe one could consider IEEE’s intepretation of “not a number” to be more like “not 1 particular number”.

Adherence to IEEE 754 isn’t absolute. For another example, “squareRoot if the operand is less than zero” is listed as an operation that should return a quiet NaN by default. Julia throws a `DomainError`, Python’s `math.sqrt` raises a `ValueError`, and NumPy’s `sqrt` prints a `RuntimeWarning` the first time and returns `nan`.

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 19, 2022, 3:44am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/11 "2022-07-19T03:44:47Z")

</div>

All this said, if your data are naturally all floating point numbers and for any reason you _want_ to use `NaN` like one would normally use `missing`, you certainly _can_ do so (e.g. [GitHub - brenhinkeller/NaNStatistics.jl: Fast summary statistics, histograms, and binning – ignoring NaNs](https://github.com/brenhinkeller/NaNStatistics.jl)) – in that particular case it has some potential performance benefits (doesn’t require an extra mask, can be SIMD’d more easily, etc.).

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [July 19, 2022, 10:38am UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/12 "2022-07-19T10:38:43Z")

</div>

Makes me wonder if `Union{<:AbstractFloat, Missing}` elements can be implemented with NaNs instead of a `UInt8` type tags somehow. However, a particular concrete type not sharing the structure of the parametric supertype would be a weird outlier, and I can’t imagine how to make indexing a `missing` element from a float array consistent with doing so from any other array. NaNStatistics.jl doesn’t deal with `missing`s at all, so it gets to use `Aᵢ==Aᵢ` as a `!isnan` for all `Number`s and to skip float NaNs like `missing`s (although, I think I can break this with a custom non-float `Number` where `==` is weirdly `false`).

---

<div class="post-metadata">

**Author:** ![brenhinkeller](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brenhinkeller/32/36648_2.png) [@brenhinkeller](https://discourse.julialang.org/u/brenhinkeller)\
**Post date:** [July 19, 2022, 5:59pm UTC](https://discourse.julialang.org/t/is-there-any-reason-to-use-nan-instead-of-missing/84396/13 "2022-07-19T17:59:27Z")

</div>

In principle you certainly could, because there’s something like 48 bits of payload in Float64 NaNs which can take arbitrary values – and apparently people do all sorts of wild things with this. As I learned from this thread [Slack](https://julialang.slack.com/archives/C680MM7D4/p1657867380480259) about [tinylisp](https://github.com/Robert-van-Engelen/tinylisp), modern [x86 64-bit pointers only actually use 48 bits](https://en.wikipedia.org/wiki/X86-64#Virtual_address_space_details), so you can fit a whole pointer in the payload of a NaN (which is apparently called “NaN boxing”)
