# Converting missing type into float

**URL:** <https://discourse.julialang.org/t/converting-missing-type-into-float/7740>\
**Category:** Data\
**Tags:** data, type\
**Created:** [December 13, 2017, 3:01pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740 "2017-12-13T15:01:53Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fred](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fred/32/14175_2.png) [@Fred](https://discourse.julialang.org/u/Fred)\
**Post date:** [December 13, 2017, 3:01pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/1 "2017-12-13T15:01:53Z")

</div>

Hi,

The new “missing” type introduced in DataFrames is a source of many errors in all the Julia programs I have made previously and I would like to find a way to remove this missing type.

For example the competerank() function from package MultipleTesting is not compatible with missing type.

the command  
`r = competerank(pv)`  
now gives and error :  
MethodError: no method matching competerank(::Array{Union{Float64, Missings.Missing},1})

I tried to replace it by :  
`r = competerank(skipmissing(pv))`

or

```julia
pv = convert(Array,pv)
r = competerank(pv)

```

without success 🙄  
Thank you for your help !

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [December 13, 2017, 3:09pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/2 "2017-12-13T15:09:40Z")

</div>

I’m not quite sure what’s being asked here, but I suggest that if your dataset _really does_ contain fields of missing data that you define the appropriate functions for the `Missing` type. After all, that’s the whole reason why `Missing` exists.

Note that `skipmissing` returns an iterator, not an array. If you want an array you can do `collect(skipmissing(pv))`. If you want to replace the missing values (you seemed to imply that you’d want to replace them with floats) you can do `Missings.replace(pv, x)`, but be aware that `replace` also returns an iterator so if you want an array you again have to do `collect`.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 13, 2017, 3:16pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/3 "2017-12-13T15:16:23Z")

</div>

It’s too bad that `competerank` doesn’t accept any iterable, but that limitation comes from `sortperm`. I guess it would be possible to write a version which automatically collects into an array for convenience. Could you file an issue against StatsBase?

Note that this is not special to `missing`, it also occurred with `NA` AFAICT.

---

<div class="post-metadata">

**Author:** ![Fred](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fred/32/14175_2.png) [@Fred](https://discourse.julialang.org/u/Fred)\
**Post date:** [December 13, 2017, 3:34pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/4 "2017-12-13T15:34:05Z")

</div>

thank you @ExpandingMan, @nalimilan

with Julia 0.5 and corresponding version of DataFrames I could read my data into a DataFrame and then use directly a column of the table (pv in my example) in statistical functions such as competerank()

I did a quick test following the suggestion of @ExpandingMan, and it seems that the compatibility is restored using this line

`pv = collect(skipmissing(pv))`

at least for the functions

```julia
competerank()
adjust()

```

there is may be a slowdown but my priority is to keep my old programs working. Thanks !

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [December 13, 2017, 3:36pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/5 "2017-12-13T15:36:50Z")

</div>

You should find that for the most part, with the latest version of dataframes you can again “use directly a column of the table”, but some functions simply don’t make sense in the presence of `missing`, and their behavior must be defined in these cases.

I would be interested to know what the philosophy of `StatsBase` is toward `Missing`. Is lifting considered to be a default behavior?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 13, 2017, 3:45pm UTC](https://discourse.julialang.org/t/converting-missing-type-into-float/7740/6 "2017-12-13T15:45:14Z")

</div>

> [@Fred](#):
>
> with Julia 0.5 and corresponding version of DataFrames I could read my data into a DataFrame and then use directly a column of the table (pv in my example) in statistical functions such as competerank()

I see, this was actually possible because `DataArray` lied about its element type and pretended it could not contain missing values. That situation had to be fixed: functions which support missing values now have to opt-in to accept such arrays, or the user needs to remove missing values manually.

> [@ExpandingMan](#):
>
> I would be interested to know what the philosophy of StatsBase is toward Missing. Is lifting considered to be a default behavior?

At this point the philosophy is that you need to skip missing values manually before calling StatsBase functions. I guess some functions could accept missing values and return `missing` if they find one, but that’s not terribly useful so that’s low priority.
