# Nanmean options?

**URL:** <https://discourse.julialang.org/t/nanmean-options/4994>\
**Category:** General Usage\
**Created:** [July 21, 2017, 8:09am UTC](https://discourse.julialang.org/t/nanmean-options/4994 "2017-07-21T08:09:38Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Daneel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/daneel/32/175_2.png) [@Daneel](https://discourse.julialang.org/u/Daneel)\
**Post date:** [July 21, 2017, 8:09am UTC](https://discourse.julialang.org/t/nanmean-options/4994/1 "2017-07-21T08:09:38Z")

</div>

[Up until Julia 0.6 `mean` ignored NaN’s](https://github.com/JuliaLang/julia/pull/12563). What are the options for those of us that want a simple `nanmean` function?

Somewhere I remember seeing an alternative function in Base but I haven’t been able to find it again.

There were two options mentioned in the commit linked above: NanMath.jl and DataArrays.jl. NaNMath looks like a good option if there’s nothing available in Base but DataArrays looks like overkill for a simple nanmean function.

> **[GitHub - JuliaMath/NaNMath.jl: Julia math built-ins which return NaN and...](https://github.com/JuliaMath/NaNMath.jl)**
>
> Julia math built-ins which return NaN and accumulator functions which ignore NaN - GitHub - JuliaMath/NaNMath.jl: Julia math built-ins which return NaN and accumulator functions which ignore NaN

> **[GitHub - JuliaStats/DataArrays.jl: DEPRECATED: Data structures that allow...](https://github.com/JuliaStats/DataArrays.jl)**
>
> DEPRECATED: Data structures that allow missing values - GitHub - JuliaStats/DataArrays.jl: DEPRECATED: Data structures that allow missing values

Matlab (just for those searching for nanmean usage like in Matlab)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 21, 2017, 2:03pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/2 "2017-07-21T14:03:18Z")

</div>

You can define a function, but

```julia
mean(filter(!isnan, x))

```

is already very compact.

---

<div class="post-metadata">

**Author:** ![bobportmann](https://avatars.discourse-cdn.com/v4/letter/b/a88e57/32.png) [@bobportmann](https://discourse.julialang.org/u/bobportmann)\
**Post date:** [July 21, 2017, 2:32pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/3 "2017-07-21T14:32:22Z")

</div>

Yes, but that doesn’t work if x is more that 1D and you want to apply mean to one dimension, e. g.

```julia
mean(filter(!isnan, x), 2)

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 21, 2017, 4:34pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/4 "2017-07-21T16:34:53Z")

</div>

Define `nanmean` as I suggested, then `nanmean(X, d, dims)` with `mapslices`.

---

<div class="post-metadata">

**Author:** ![bobportmann](https://avatars.discourse-cdn.com/v4/letter/b/a88e57/32.png) [@bobportmann](https://discourse.julialang.org/u/bobportmann)\
**Post date:** [July 21, 2017, 6:07pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/5 "2017-07-21T18:07:03Z")

</div>

OK. It would be nice if this was done in `NaNMath` so it was a full replacement for the functions it defines.

---

<div class="post-metadata">

**Author:** ![Daneel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/daneel/32/175_2.png) [@Daneel](https://discourse.julialang.org/u/Daneel)\
**Post date:** [July 26, 2017, 11:45am UTC](https://discourse.julialang.org/t/nanmean-options/4994/6 "2017-07-26T11:45:49Z")

</div>

Here is a concise example based on @Tamas_Papp 's suggestions.

```julia
nanmean(x) = mean(filter(!isnan,x))
nanmean(x,y) = mapslices(nanmean,x,y)

```

* * *

```julia
julia> y = [NaN 2 3 4;5 6 NaN 8;9 10 11 12]
3×4 Array{Float64,2}:
 NaN 2.0 3.0 4.0
   5.0 6.0 NaN 8.0
   9.0 10.0 11.0 12.0

julia> nanmean(y)
7.0

julia> nanmean(y,1)
1×4 Array{Float64,2}:
 7.0 6.0 7.0 8.0

```

---

<div class="post-metadata">

**Author:** ![deszoeke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deszoeke/32/2174_2.png) [@deszoeke](https://discourse.julialang.org/u/deszoeke)\
**Post date:** [March 5, 2018, 6:36pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/7 "2018-03-05T18:36:06Z")

</div>

Is `mean(filter(!isnan, x)` fast compared to filtering inside the accumulator? Does it require more memory?

`mapslices` solution is awesome but seems obscure to me, compared to `nanmean`. I always need `nanmean` for processing geophysical data, and would find it useful in StatsBase.

Also note, like Matlab’s, this `nanmean` along multiple dimensions is not associative in general, but depends on order of operations on dimensions,  
`nanmean(nanmean((x,1),2) != nanmean(nanmean((x,2),1)`.

---

<div class="post-metadata">

**Author:** ![deszoeke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deszoeke/32/2174_2.png) [@deszoeke](https://discourse.julialang.org/u/deszoeke)\
**Post date:** [May 18, 2018, 5:27pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/8 "2018-05-18T17:27:10Z")

</div>

Some useful comments are posted after the closed issue [Fail to ignore NaN when calculating mean of an array #4552](https://github.com/JuliaLang/julia/issues/4552#issuecomment-390272169). Note the appropriate place for comments is here in Discourse. So to move the thread here:

@timholy’s Images.jl and `meanfinite` have fast, flexible dimensioned averages that ignore non-finite values.

I have reworked `meanfinite` as `condmean` in [ConditionalMean](https://github.com/deszoeke/ConditionalMean.jl) to accumulate subject to an arbitrary condition (which could be to ignore sentinel values, i.e. `NaN`, -999). It also can average a callable function of the array. `nanmean()` and `nanstd()` methods are provided. Tests and pull requests are welcome.

Now this devolves into a ease-of-use complaint: The unsettledness of standard(s) for how to handle missing data is an impediment to developing useful functions and methods. The long list of breaking, late-breaking, and deprecated ways to do this include `NA`, `missing`, `Null`, Unions thereof, and sentinel values for numeric types (e.g. `NaN`).

This diversity and changing architecture leads to severe usability problems. Searching the discussions, I find different computer science and data science philosophies and practical reasons for one approach over another. I respect these arguments, but it’s ~~impossible~~ hard to tell what works, what’s supported, deprecated, or broken.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 18, 2018, 5:40pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/9 "2018-05-18T17:40:32Z")

</div>

> [@deszoeke](#):
>
> The unsettledness of standard(s) for how to handle missing data is an impediment to developing useful functions and methods.

I sympathize with the frustration about this issue, but note that things are in transition now because of the anticipated efficiency gain for small unions in `v0.7`. Hopefully the semantics will settle soon after the next stable release.

---

<div class="post-metadata">

**Author:** ![bjarthur](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bjarthur/32/9638_2.png) [@bjarthur](https://discourse.julialang.org/u/bjarthur)\
**Post date:** [June 8, 2020, 8:17pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/10 "2020-06-08T20:17:25Z")

</div>

thanks daneel. here is a more generic version of your code above:

```julia
using Statistics, Test

_nanfunc(f, A, ::Colon) = f(filter(!isnan, A))
_nanfunc(f, A, dims) = mapslices(a->_nanfunc(f,a,:), A, dims=dims)
nanfunc(f, A; dims=:) = _nanfunc(f, A, dims)

A = [1 2 3; 4 5 6; 7 8 9; NaN 11 12]

@test isapprox(nanfunc(mean, A), mean(filter(!isnan, A)))
@test nanfunc(mean, A, dims=1) == [4.0 6.5 7.5]
@test nanfunc(mean, A, dims=2) == transpose([2.0 5.0 8.0 11.5])

@test isapprox(nanfunc(var, A), var(filter(!isnan, A)))
@test nanfunc(var, A, dims=1) == [9.0 15.0 15.0]
@test nanfunc(var, A, dims=2) == transpose([1.0 1.0 1.0 0.5])

```

---

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [April 8, 2021, 7:31am UTC](https://discourse.julialang.org/t/nanmean-options/4994/11 "2021-04-08T07:31:14Z")

</div>

nanmean(y,1) throws an error for me:

MethodError: no method matching mapslices(::typeof(nanmean), ::Array{Float64,2}, ::Int64)

---

<div class="post-metadata">

**Author:** ![Daneel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/daneel/32/175_2.png) [@Daneel](https://discourse.julialang.org/u/Daneel)\
**Post date:** [April 8, 2021, 8:20am UTC](https://discourse.julialang.org/t/nanmean-options/4994/12 "2021-04-08T08:20:36Z")

</div>

Good catch. At some point `mapslices` changed how it inputs dimensions. Apparently I never use that case in my code so I hadn’t noticed it. Fortunately the fix is simple, `y` ⟹ `dims=y`

```julia
nanmean(x) = mean(filter(!isnan,x))
nanmean(x,y) = mapslices(nanmean,x,dims=y)

```

---

<div class="post-metadata">

**Author:** ![thompsonmj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thompsonmj/32/25172_2.png) [@thompsonmj](https://discourse.julialang.org/u/thompsonmj)\
**Post date:** [June 3, 2021, 7:56pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/14 "2021-06-03T19:56:05Z")

</div>

When using `mean(filter(!isnan, A), dims=2)` for a 2D matrix A, the filter flattens it into a vector before taking the mean. How can I do a mean over the columns of matrix ignoring the NaN values?

---

<div class="post-metadata">

**Author:** ![aramirezreyes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aramirezreyes/32/42573_2.png) [@aramirezreyes](https://discourse.julialang.org/u/aramirezreyes)\
**Post date:** [June 3, 2021, 8:33pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/15 "2021-06-03T20:33:26Z")

</div>

Maybe [GitHub - brenhinkeller/NaNStatistics.jl: Fast summary statistics, histograms, and binning – ignoring NaNs](https://github.com/brenhinkeller/NaNStatistics.jl) could help here.

---

<div class="post-metadata">

**Author:** ![thompsonmj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thompsonmj/32/25172_2.png) [@thompsonmj](https://discourse.julialang.org/u/thompsonmj)\
**Post date:** [June 3, 2021, 8:57pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/16 "2021-06-03T20:57:14Z")

</div>

Cool! That worked for me, thanks! Seems like this should be achievable with the `skipmissing` feature in [Statistics](https://docs.julialang.org/en/v1/stdlib/Statistics/), but it doesn’t seem to accept both `skipmissing` and `dims`.

---

<div class="post-metadata">

**Author:** ![ZekeMarshall](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zekemarshall/32/206876_2.png) [@ZekeMarshall](https://discourse.julialang.org/u/ZekeMarshall)\
**Post date:** [February 11, 2024, 10:55am UTC](https://discourse.julialang.org/t/nanmean-options/4994/17 "2024-02-11T10:55:01Z")

</div>

Thanks, this thread is a lifesaver and such a feature does feel like a conspicuous absence! Particularly the lack of support for using both `skipmissing` and `dims=` as @thompsonmj mentions.

In my case I needed to find the minimum value in each column of a 2d array `mat`, ignoring zeros. I find that all three of the methods below work:

`mat = rand(0:5, 5, 11)`

1: `mapslices(x -> minimum(skipmissing(x)), replace(mat, 0 => missing), dims = 1)`  
2: `mapslices(x -> minimum(filter(!isnan, x)), replace(mat, 0 => NaN), dims = 1)`  
3: `mapslices(x -> minimum(filter(!iszero, x)), mat, dims = 1)`

However, these approaches only work if there are no columns which only contain zeros. Which will lead to the error:

`ERROR: MethodError: reducing over an empty collection is not allowed`

There is a huge performance penalty when doing this with anonymous functions, e.g.

`@elapsed minimum(mat, dims = 1)` ~ 3.8735e-5s  
`@elapsed mapslices(x -> minimum(filter(!iszero, x)), mat, dims = 1)` ~ 0.159694082s

Defining the function beforehand helps, e.g.

`nzmin(x) = minimum(filter(!iszero, x))`  
`nzmin(x,y) = mapslices(nzmin, mat, dims = y)`  
`@elapsed nzmin(mat,1)` ~ 0.000339807s

or using a modified version of the generalised function @bjarthur outlines:

`_nzfunc(f, A, ::Colon) = f(filter(!iszero, A))`  
`_nzfunc(f, A, dims) = mapslices(a->_nzfunc(f,a,:), A, dims=dims)`  
`nzfunc(f, A; dims=:) = _nzfunc(f, A, dims)`

`@elapsed nzfunc(minimum, mat, dims=1)` ~ 0.000850458

Which incurs a small penalty in passing a function through, relative to defining a specific function.

I’m new new to Julia, so may have missed something here!

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [February 11, 2024, 1:12pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/18 "2024-02-11T13:12:58Z")

</div>

1. `minimum(filter(!iszero, x))` materializes an actual filtered array in memory, making it not that efficient. The `skip` function (from `Skipper.jl`) would be better here.

2. `mapslices` seem to be suboptimal, regular `map()` on `eachcol(mat)` can be faster.

3. For measuring time taken by fast pieces of code (\<\< 1 sec), use `BenchmarkTools.jl` instead of `@time`/`@elapsed`.

Comparison:

```julia
# baseline:
julia> @btime minimum($mat, dims = 1)
  215.762 ns (3 allocations: 240 bytes)

julia> using Skipper

# skip zeros using mapslices
julia> @btime mapslices(x -> minimum(skip(iszero, x)), $mat, dims=1)
  499.139 ns (15 allocations: 624 bytes)

# skip zeros using map(eachcol)
julia> @btime map(x -> minimum(skip(iszero, x)), eachcol($mat))
  64.626 ns (1 allocation: 144 bytes)

```

---

<div class="post-metadata">

**Author:** ![ZekeMarshall](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zekemarshall/32/206876_2.png) [@ZekeMarshall](https://discourse.julialang.org/u/ZekeMarshall)\
**Post date:** [February 11, 2024, 6:50pm UTC](https://discourse.julialang.org/t/nanmean-options/4994/19 "2024-02-11T18:50:40Z")

</div>

Thank you very much @aplavin!! I didn’t know about Skipper, just what I needed!
