# Filter of skipmissing fails

**URL:** <https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249>\
**Category:** General Usage\
**Created:** [February 27, 2019, 6:54am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249 "2019-02-27T06:54:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [February 27, 2019, 6:54am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/1 "2019-02-27T06:54:18Z")

</div>

MWE:

```julia
julia> x = [missing, 1, 2, -1, -2];
julia> filter(e->(e>0), skipmissing(x))
ERROR: MethodError: no method matching filter(::getfield(Main, Symbol("##23#24")), ::Base.SkipMissing{Array{Union{Missing, Int64},1}})
julia> filter(e->(e>0), collect(skipmissing(x)))
2-element Array{Int64,1}:
 1
 2

```

Wouldn’t it be desirable to have the first filter call work? Potentially it could amount to the second one though it’s probably not ideal as per small benchmark below, likely better can be done:

```julia
julia> λ(e) = (e>0);
julia> filter_skipmissing(λ, x) = filter(e->(!ismissing(e) && λ(e)), x)
julia> x = [(rand()<0.3) ? missing : randn() for i = 1:500_000];
julia> @btime filter($λ, collect(skipmissing($x)));
  8.071 ms (38 allocations: 8.00 MiB)
julia> @btime filter_skipmissing($λ, $x);
  5.984 ms (18 allocations: 3.38 MiB)

```

I guess more generally something like this could be used:

```julia
julia> filter_itr(λ, itr) = [e for e ∈ itr if λ(e)];
julia> @btime filter_itr($λ, skipmissing($x));
  5.672 ms (22 allocations: 3.00 MiB)

```

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [February 27, 2019, 7:40am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/2 "2019-02-27T07:40:08Z")

</div>

Good question!

`filter` is for filtering arrays. Since `skipmissing` is an iterable, you can instead use `Iterators.filter`. An advantage of this approach is that it keeps the lazy property of the iterable. And if you do need to collect the output, call `collect` on the result:

```julia
julia> x = [missing, 1, 2, -1, -2];

julia> itr = Iterators.filter(e -> e > 0, skipmissing(x))
Base.Iterators.Filter{getfield(Main, Symbol("##22#23")),Base.SkipMissing{Array{Union{Missing, Int64},1}}}(getfield(Main, Symbol("##22#23"))(), Base.SkipMissing{Array{Union{Missing, Int64},1}}(Union{Missing, Int64}[missing, 1, 2, -1, -2]))

julia> collect(itr)
2-element Array{Int64,1}:
 1
 2

```

When collecting in the end, performance is about the same as your fastest version:

```julia
julia> λ(e) = (e>0);

julia> x = [(rand()<0.3) ? missing : randn() for i = 1:500_000];

julia> filter_itr(λ, itr) = [e for e ∈ itr if λ(e)];

julia> @btime filter_itr($λ, skipmissing($x)); # your fastest version
  6.010 ms (22 allocations: 3.00 MiB)

julia> @btime collect(Iterators.filter($λ, skipmissing($x)));
  5.959 ms (20 allocations: 3.00 MiB)

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [February 27, 2019, 7:49am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/3 "2019-02-27T07:49:10Z")

</div>

> [@bennedich](#):
>
> `filter` is for filtering arrays. Since `skipmissing` is an iterable, you can instead use `Iterators.filter` .

I don’t think that there is anything that in principle would preclude `Base.filter` from working on iterables.

IMO the distinction between `Iterators.*` and `Base.*` methods should be that the latter materialize the result while the former return an iterator.

---

<div class="post-metadata">

**Author:** ![tlienart](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlienart/32/7640_2.png) [@tlienart](https://discourse.julialang.org/u/tlienart)\
**Post date:** [February 27, 2019, 10:30am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/4 "2019-02-27T10:30:26Z")

</div>

Ok thanks that’s great and I see the ambiguity in trying to provide a default: should it be the iterator or should it be the materalised version.

I’m unfamiliar with `Iterators` so that will be motivation to look into it a bit more!

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [February 27, 2019, 11:02am UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/5 "2019-02-27T11:02:51Z")

</div>

As a side note, I hardly find `Base.filter` useful in my use cases. Is there a good example that motivates the existence of non-lazy versions in `Base` or it is just user-friendliness to avoid explicit calls to `collect`?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [February 27, 2019, 1:21pm UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/6 "2019-02-27T13:21:09Z")

</div>

I think the interface accumulated historically, also there are some tricky questions (eg should `filter` narrow type like `collect` — currently it doesn’t, and it would be come type unstable).

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 27, 2019, 3:59pm UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/7 "2019-02-27T15:59:49Z")

</div>

This is also going to cause problems if you initially write code for a subset of data that doesn’t have missing values. Then when you realize the data actually does have missing values, you can’t just sprinkle `skipmissing` to all your function calls.

Edit: I posted [an issue](https://github.com/JuliaLang/julia/issues/31188) on github about this.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [February 27, 2019, 7:44pm UTC](https://discourse.julialang.org/t/filter-of-skipmissing-fails/21249/8 "2019-02-27T19:44:42Z")

</div>

> [@juliohm](#):
>
> As a side note, I hardly find `Base.filter` useful in my use cases. Is there a good example that motivates the existence of non-lazy versions in `Base`

In addition to what Tamas mentioned, there’s also `filter!` which can be very useful. It’d be confusing I think if `filter` returned an iterable while `filter!` modified an array.
