# Extending \`dropmissing\` to AbstractArrays

**URL:** <https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753>\
**Category:** General Usage\
**Tags:** question, arrays, missing-values\
**Created:** [November 26, 2023, 4:40pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753 "2023-11-26T16:40:24Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 26, 2023, 4:40pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753/1 "2023-11-26T16:40:24Z")

</div>

From another thread, there was discussion of `skipmissing`, and it prompted me to think `dropmissing` is a convenient way to get rid of those pesky `missing`s. DataFrames has `dropmissing` but regular Arrays don’t. How could it be extended? And for higher order tensors? This is the question.

Here is an initial suggestion:

```julia
julia> using Random, Missings

julia> Random.seed!(345);

julia> q = [rand() < 0.2 ? missing : rand() for i in 1:6,j in 1:6]
6×6 Matrix{Union{Missing, Float64}}:
 0.956056 0.567035 0.401763 0.457714 0.361413 0.55555
 0.810208 0.595364 0.734469 0.591841 0.75449 0.241072
  missing 0.651234 0.683582 0.362204 0.908998 0.811716
  missing 0.814048 0.752948 0.194032 0.32284 missing
 0.336461 0.167354 0.591084 0.37456 0.0637121 0.16276
 0.574073 0.363261 0.59856 0.0997761 0.364427 0.695076

julia> foldl((r,k)->stack(filter(s->!any(ismissing.(s)), eachslice(r, dims=k))), 1:ndims(q); init=q)
6×4 Matrix{Union{Missing, Float64}}:
 0.956056 0.810208 0.336461 0.574073
 0.567035 0.595364 0.167354 0.363261
 0.401763 0.734469 0.591084 0.59856
 0.457714 0.591841 0.37456 0.0997761
 0.361413 0.75449 0.0637121 0.364427
 0.55555 0.241072 0.16276 0.695076

```

The `foldl` removes any ‘slice’ which contains a `missing`. But, it’s **wrong** , since it removes all `missing`s on the first round, and depends on the order of dimensions. How can this be fixed?

ADDED:

For example, this is more correct, but still might be wasteful in case of very tall or wide matrices and somewhat allocating:

```julia
julia> nonmissingtype(eltype(q)).(
  q[map(i->[!any(ismissing,s) for s in eachslice(q, dims=i)], 
  1:ndims(q))...])
4×4 Matrix{Float64}:
 0.567035 0.401763 0.457714 0.361413
 0.595364 0.734469 0.591841 0.75449
 0.167354 0.591084 0.37456 0.0637121
 0.363261 0.59856 0.0997761 0.364427

```

This is less allocating, using a view (which might not be good enough for many downstream usage):

```julia
@view q[map(i->[!any(ismissing,s) for s in eachslice(q, dims=i)], 1:ndims(q))...]

```

Now the problem is speed, as the comprehension reads the matrix `ndims` times, instead of once.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [November 26, 2023, 6:00pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753/2 "2023-11-26T18:00:30Z")

</div>

You might be looking for `Filter(!ismissing)` (using Transducers.jl).

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 26, 2023, 6:06pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753/3 "2023-11-26T18:06:45Z")

</div>

Don’t think so:

> [@ParadaCarleton](#):
>
> Filter(!ismissing)

doesn’t keep the dimensionality of a tensor (which requires dropping some non-missing elements too).

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [November 26, 2023, 8:52pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753/5 "2023-11-26T20:52:14Z")

</div>

> [@Dan](#):
>
> `q[map(i->[!any(ismissing,s) for s in eachslice(q, dims=i)], 1:ndims(q))...]`

One alternative that seems to allocate a bit less:

```julia
q[ntuple(i->[!any(ismissing,s) for s in eachslice(q, dims=i)], ndims(q))...]

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 26, 2023, 9:53pm UTC](https://discourse.julialang.org/t/extending-dropmissing-to-abstractarrays/106753/6 "2023-11-26T21:53:15Z")

</div>

Yes. `ntuple` is better than a `Vector`. But the crux would be to scan the source Array (tensor) just once to find all the relevant subsets. I’ll try working it out again…

A FEW MINS LATER:  
This is a function which goes through tensor once:

```julia
function tensordropmissing(Q::AbstractArray{T,N}) where {T,N}
    s = size(Q)
    goodidxs = ntuple(i->trues(s[i]),N)
    for I in CartesianIndices(Q)
        if ismissing(Q[I])
            for k in 1:N
                goodidxs[k][I[k]] = false
            end
        end
    end
    return @view Q[goodidxs...]
end

```

With it, and `q` as in OP:

```julia
julia> tensordropmissing(q)
4×4 view(::Matrix{Union{Missing, Float64}}, [1, 2, 5, 6], [2, 3, 4, 5]) with eltype Union{Missing, Float64}:
 0.567035 0.401763 0.457714 0.361413
 0.595364 0.734469 0.591841 0.75449
 0.167354 0.591084 0.37456 0.0637121
 0.363261 0.59856 0.0997761 0.364427

julia> nonmissingtype(eltype(q)).(tensordropmissing(q))
4×4 Matrix{Float64}:
 0.567035 0.401763 0.457714 0.361413
 0.595364 0.734469 0.591841 0.75449
 0.167354 0.591084 0.37456 0.0637121
 0.363261 0.59856 0.0997761 0.364427

```

and almost no allocation:

```julia
julia> @btime tensordropmissing($q);
  187.650 ns (6 allocations: 384 bytes)

```
