# Improve type handling in skipmissing

**URL:** <https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513>\
**Category:** Internals & Design\
**Created:** [January 18, 2020, 8:05am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513 "2020-01-18T08:05:06Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alex1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alex1/32/11771_2.png) [@Alex1](https://discourse.julialang.org/u/Alex1)\
**Post date:** [January 18, 2020, 8:05am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/1 "2020-01-18T08:05:06Z")

</div>

Would be nice to improve `skipmissing`, maybe always apply `collect`? If the reason not to apply it is the performance, I would say - predictability and principle of least surprise is more important than performance.

```julia
serie = [1.0, missing, 1.0, 4.0]
function process(serie)
  [value for value in serie]
end

println(process(serie) == [1.0, missing, 1.0, 4.0])
# => missing - Good!

println(skipmissing(process(serie)) == skipmissing([1.0, missing, 1.0, 4.0]))
# => false - Surprise, not good.

println(collect(skipmissing(process(serie))) == collect(skipmissing([1.0, missing, 1.0, 4.0])))
# => true

```

---

<div class="post-metadata">

**Author:** ![Sijun](https://avatars.discourse-cdn.com/v4/letter/s/b2d939/32.png) [@Sijun](https://discourse.julialang.org/u/Sijun)\
**Post date:** [January 18, 2020, 8:38am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/2 "2020-01-18T08:38:55Z")

</div>

Once we know that skipmissing is a lazy evaluation, I think it should not come as a surprise that we can’t know the result until it is evaluated.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [January 18, 2020, 8:50am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/3 "2020-01-18T08:50:02Z")

</div>

The reason for `skipmissing` to return an iterator instead of a collected array is not only performance but also space. Think about it, when you’re handling a large dataset (maybe a 10GB csv) and you just always collect after skipmissing, you’re going to maybe double the memory requirements (at least in the short term) because a new array for the non-missing entries has to be created, instead of using the existing memory and just… skipping missings when iterating.

---

<div class="post-metadata">

**Author:** ![cstjean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cstjean/32/1444_2.png) [@cstjean](https://discourse.julialang.org/u/cstjean)\
**Post date:** [January 18, 2020, 9:27am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/4 "2020-01-18T09:27:33Z")

</div>

```julia
julia> (x for x in serie) == (x for x in serie)
false

```

It seems that all iterators are compared strictly for identity (`===`), which I find a bit surprising, but I suppose there was a good reason for it.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [January 18, 2020, 9:44am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/5 "2020-01-18T09:44:00Z")

</div>

Note that you’re comparing generators in your example, not iterators. Iterators only exist as part of an implemented interface of functions, not as concrete subtypes of some abstract type. See [here](https://docs.julialang.org/en/v1/manual/interfaces/#man-interface-iteration-1). Collecting a generator or an iterator to check whether they produce the same output goes against their core idea of lazy evaluation, though I suppose their fields could be checked for identity instead.

In any case, checking for identity is defined as a fallback for any comparison in case no specialized version exists. You can check for yourself via `julia> @edit (x for x in serie) == (x for x in serie)` and you’ll see that you will end up in the generic definition of `==(x,y)` in `operators.jl`.

---

<div class="post-metadata">

**Author:** ![cstjean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cstjean/32/1444_2.png) [@cstjean](https://discourse.julialang.org/u/cstjean)\
**Post date:** [January 18, 2020, 9:56am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/6 "2020-01-18T09:56:24Z")

</div>

> [@Sukera](#):
>
> Collecting a generator or an iterator to check whether they produce the same output goes against their core idea of lazy evaluation

Does it? I would believe that `===` and `IdDict`s are perfectly valid ways of comparing iterators by their identity. Using identity for `==` might be correct, but it seems like a lost opportunity for a meaningful and distinct operation. How else should you check that the two iterators contain the same data, without a wasteful `collect`? `all(==, iter1, iter2)` doesn’t work, sadly. `all(Iterators.map(==, iter1, iter2))` ?

FWIW Python3 works the same way, so I must be missing something.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [January 18, 2020, 10:10am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/7 "2020-01-18T10:10:15Z")

</div>

Iterating over iterators to check whether they produce the same result is fine, _collecting_ (i.e., allocating space to hold the resulting collection) to check for equality by default is not since you don’t necessarily know how large the collection is when comparing.

Also:

> though I suppose their fields **[of the struct used to dispatch on the iteration interface]** could be checked for identity instead.

So in that sense, I think we agree 🙂 I do wonder if you really want to iterate over them, since `IteratorSize`, `length` and `size` are only optional for iterators in general and thus comparing iterators by iterating can potentially lead to very long unintended runtimes.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [January 18, 2020, 7:17pm UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/8 "2020-01-18T19:17:03Z")

</div>

> [@Alex1](#):
>
> Would be nice to improve `skipmissing` , maybe always apply `collect` ? If the reason not to apply it is the performance, I would say - predictability and principle of least surprise is more important than performance.

If that is the behaviour you desire and you don’t find yourself caring about performance / memory allocations, you can simply define

```julia
myskipmissing(args...) = collect(skipmissing(args...))

```

The reason I think the current default is better is that you can always collect an iterator, but once one has already been materialized for you, there’s no undoing those allocations.

The reason Julia doesn’t provide both is that we want to avoid cluttering the namespace of functions exported from base, so we mostly avoid providing functions that can be trivially written as the composition of two functions (I.e. `collect` and `skipmissing`).

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [January 19, 2020, 3:08am UTC](https://discourse.julialang.org/t/improve-type-handling-in-skipmissing/33513/9 "2020-01-19T03:08:19Z")

</div>

> [@Sukera](#):
>
> comparing iterators by iterating can potentially lead to very long unintended runtimes

It’s not just performance. Iterators may be stateful (e.g., `eachline`) so `==` cannot iterate over elements when the input is `(x for x in stateful_iter)` or `skipmissing(stateful_iter)`.
