# Ignoring NaN in elementwise aggregations

**URL:** <https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431>\
**Category:** General Usage\
**Tags:** question\
**Created:** [December 18, 2019, 9:04am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431 "2019-12-18T09:04:57Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 18, 2019, 9:04am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/1 "2019-12-18T09:04:57Z")

</div>

I would like to do calculations like

```julia
mean(a_bunch_of_similar_arrays)

```

but with `NaN`s quietly ignored.

MWE with an inefficient and inelegant solution:

```julia
A1 = [NaN 1; 2 3]

A2 = [1 NaN; 2 3]

As = [A1, A2]

function elementwise_ignoreNaN(f, arrays)
    ax = axes(first(arrays))
    I = CartesianIndices(ax)
    @assert all(a -> axes(a) == ax, arrays)
    map(I) do ι
        x = filter(!isnan, [a[ι] for a in arrays])
        f(x)
    end
end

elementwise_ignoreNaN(mean, As) # [1 1; 2 3]

```

but I wonder if this functionality exists already in some package, or if I could do this better.

---

<div class="post-metadata">

**Author:** ![tkluck](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkluck/32/15769_2.png) [@tkluck](https://discourse.julialang.org/u/tkluck)\
**Post date:** [December 18, 2019, 10:12am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/2 "2019-12-18T10:12:21Z")

</div>

Any reason not to use broadcasting? E.g. something similar to

```julia
julia> nanless(f, a...) = f(filter(!isnan, a))
nanless (generic function with 1 method)

julia> nanless.(mean, A1, A2)
2×2 Array{Float64,2}:
 1.0 1.0
 2.0 3.0

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 18, 2019, 10:20am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/3 "2019-12-18T10:20:35Z")

</div>

I have about 5000 such arrays (sorry, should have said), and I am not sure that broadcasting is the good match for that.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [December 18, 2019, 10:52am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/4 "2019-12-18T10:52:27Z")

</div>

Could [`skipmissing`](https://docs.julialang.org/en/v1/manual/missing/#Skipping-Missing-Values-1) help if you use `missing` instead of `NaN`? For some reason it doesn’t seem to work on your MWE but it promises to be the efficient way to filter missing values in calculations.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [December 18, 2019, 11:45pm UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/5 "2019-12-18T23:45:11Z")

</div>

Maybe you are asking for a package that does this out-of-the-box, but I think using [`LazyArrays.@~`](https://github.com/JuliaArrays/LazyArrays.jl) to create `Broadcasted` objects is a neat approach:

```julia
xs = [[NaN 1; 2 3], [1 NaN; 2 3]]

add!!(a, b) = a .+ b
add!!(a::AbstractArray, b) = a .+= b
sumpw(f, xs) = mapreduce(f, add!!, xs)

using LazyArrays: @~
fillnan(x) = isnan(x) ? zero(x) : x
sumpw(x -> (@~ fillnan.(x)), xs) ./
    sumpw(x -> (@~ .!isnan.(x)), xs)

```

With this approach, you don’t have to allocate intermediate arrays like `fillnan.(x)`. The array allocated by materializing `fillnan.(xs[1]) +. fillnan.(xs[2])` is used for a entire base case of the reduction. You can also use `mapfoldl` instead of `mapreduce` to minimize the allocation.

(I’ll probably add `add!!` to [BangBang.jl](https://github.com/tkf/BangBang.jl) so that above code would be more generic.)

This is also easily parallelizable by `using Transducers; sumpw(f, xs) = reduce(add!!, Map(f), xs)`.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 19, 2019, 7:32pm UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/6 "2019-12-19T19:32:30Z")

</div>

Have you had a look at [NaNMath.jl](https://github.com/mlubin/NaNMath.jl)?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 20, 2019, 9:27am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/7 "2019-12-20T09:27:09Z")

</div>

> [@tbeason](#):
>
> Have you had a look at [NaNMath.jl](https://github.com/mlubin/NaNMath.jl)?

Thanks for the suggestion - NaNMath.jl should simplify the filter/reduce step. I still need a neat way to map a bunch of arrays elementwise though.

> [@nilshg](#):
>
> Could [`skipmissing`](https://docs.julialang.org/en/v1/manual/missing/#Skipping-Missing-Values-1) help if you use `missing` instead of `NaN` ?

I don’t think it is appropriate for my use case, since the objects being skipped are elements of arrays in a vector, not elements of a vector.

> [@tkf](#):
>
> I think using [`LazyArrays.@~`](https://github.com/JuliaArrays/LazyArrays.jl) to create `Broadcasted` objects is a neat approach:

It is indeed neat (once I understand it 😉), but conceptually isn’t it just

```julia
As = [[NaN 1; 2 3], [1 NaN; 2 3]]
sum(A -> map(x -> isnan(x) ? zero(x) : x, A), As) ./ sum(A -> map(!isnan, A), As)

```

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [December 20, 2019, 9:43am UTC](https://discourse.julialang.org/t/ignoring-nan-in-elementwise-aggregations/32431/8 "2019-12-20T09:43:04Z")

</div>

> [@Tamas\_Papp](#):
>
> conceptually isn’t it just

That’s right. The reason why I thought it was neat was that `@~`-based implementation allocates arrays only when necessary; i.e., thanks to the clever design of the broadcasting machinery, `fillnan.(x)` and `a .+ b` are fused even though those expressions are not syntactically in the same “dot call expression.”
