# DataFrames: How to remove rows containing NaNs when there are also missings

**URL:** <https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [May 18, 2021, 9:29pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396 "2021-05-18T21:29:43Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![florian](https://avatars.discourse-cdn.com/v4/letter/f/898d66/32.png) [@florian](https://discourse.julialang.org/u/florian)\
**Post date:** [May 18, 2021, 9:29pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/1 "2021-05-18T21:29:44Z")

</div>

Assume I have the following DataFrame and want to remove rows containing NaN:

`df = DataFrame(a=[NaN, 1.1, NaN, missing, missing], b=[1.1, 2, 3, missing, NaN], c='a':'e');`

For just one column I could do something like:  
`filter(x->(ismissing(x.a) || !isnan(x.a)), df)`

To extend this to all columns I tried to use the subset function in combination with the usual DataFrame transformation syntax, but couldn’t get it to work:  
`subset(df, :a => ByRow(x->(ismissing(x) || !isnan(x))))` (works)  
`subset(df, names(df, Union{Float64, Missing}) .=> ByRow(x->(ismissing(x) || !isnan(x))))` (doesn’t work)

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 18, 2021, 9:42pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/2 "2021-05-18T21:42:51Z")

</div>

The simplest is probably:

```julia
filter(row -> all(x -> !(x isa Number && isnan(x)), row), df)

```

You can also write:

```julia
subset(df, (names(df) .=> ByRow(x -> !(x isa Number && isnan(x))))...)

```

Note that `names(df, Union{Float64, Missing})` is not fully correct, as your column could have e.g. `Any` type and still contain `NaN`.

---

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [December 15, 2023, 2:28am UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/3 "2023-12-15T02:28:03Z")

</div>

Sorry for revive this topic. But why there is no `dropna` function like pandas? We have `dropmissing` but no `dropna`…

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 15, 2023, 10:15am UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/4 "2023-12-15T10:15:09Z")

</div>

> [@liuyxpp](#):
>
> But why there is no `dropna` function like pandas?

Because pandas in the past did not have a first class support for missing values, so it used `NaN` as a surrogate.

In DataFrames.jl by design `missing` values are properly supported, so we have `dropmissing`. In Julia `NaN` should not be used to indicate missingness.

---

<div class="post-metadata">

**Author:** ![xgdgsc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xgdgsc/32/608_2.png) [@xgdgsc](https://discourse.julialang.org/u/xgdgsc)\
**Post date:** [November 28, 2024, 10:07am UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/5 "2024-11-28T10:07:37Z")

</div>

But if you do mmap array files on disk and merge them without copy to a DataFrame you cannot use a column type with union of missing. And imagine some new user come from python with deep habit of using NaN as missing. There could still be value adding the dropna function by default.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [November 28, 2024, 11:10am UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/6 "2024-11-28T11:10:39Z")

</div>

> [@florian](#):
>
> df = DataFrame(a=[NaN, 1.1, NaN, missing, missing], b=[1.1, 2, 3, missing, NaN], c=‘a’:‘e’);

Is there a reason (or reasons) why isnan() is not defined for characters?

```julia
julia> any(isnan, df[4,:])
ERROR: MethodError: no method matching isnan(::Char)

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 28, 2024, 11:38am UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/7 "2024-11-28T11:38:02Z")

</div>

> [@rocco\_sprmnt21](#):
>
> why isnan() is not defined for characters

By its definition, `NaN` is a value of floating point representations, so it is “non-sensical” to ask a non floating point if it is `NaN`. Similarly, there is no `iszero` for Char.

Seems like a legit choice to me.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [November 28, 2024, 12:40pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/8 "2024-11-28T12:40:58Z")

</div>

mmmh … will be.  
But do you admit that it sounds strange at least from an “aesthetic” point of view that `isa(NaN, Number) = true`?

```julia
filter(row -> all(x -> !(x === NaN), row), df)

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 28, 2024, 1:12pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/9 "2024-11-28T13:12:36Z")

</div>

> [@rocco\_sprmnt21](#):
>
> isa(NaN, Number) = true

is okay, since

```julia
isa(1im, Number) == true

```

and `NaN` is a special number. You can actually get it returned by math ops:

```julia
@fastmath sqrt(-2.0) == NaN

```

Of course, other choices could be made, but this is one of them.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [November 28, 2024, 10:01pm UTC](https://discourse.julialang.org/t/dataframes-how-to-remove-rows-containing-nans-when-there-are-also-missings/61396/10 "2024-11-28T22:01:49Z")

</div>

> [@Dan](#):
>
> is okay

I just noticed that it reads like:

“Is a Not\_a\_Number a Number?”  
“yes, Not\_a\_Number is a Number”
