# DataFrame has NA, what best to do?

**URL:** <https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621>\
**Category:** General Usage\
**Created:** [December 13, 2018, 4:18am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621 "2018-12-13T04:18:30Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![LeoK987](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leok987/32/4880_2.png) [@LeoK987](https://discourse.julialang.org/u/LeoK987)\
**Post date:** [December 13, 2018, 4:18am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/1 "2018-12-13T04:18:30Z")

</div>

I read in a dataframe with CSV.read that contains some “NA” elements. CSV.read categorized those columns with NA’s as String. So it’s easy for me to remove the rows with the NA’s by comparing the string of every element with “NA”. But then the columns remain of type String. How can I convert them into Float64? Or is there a better to do this?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 13, 2018, 4:46am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/2 "2018-12-13T04:46:04Z")

</div>

```julia
x = ["1.61", "1", "2"]
x_float = parse.(Float64, x)

```

---

<div class="post-metadata">

**Author:** ![LeoK987](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leok987/32/4880_2.png) [@LeoK987](https://discourse.julialang.org/u/LeoK987)\
**Post date:** [December 13, 2018, 4:57am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/3 "2018-12-13T04:57:41Z")

</div>

Thank you. That was very simple.

I was trying to give `types=Dict(:Col1=Union{Float64, Missing})` to CSV.read, and then use `ismissing` to check every element. That seems to work, but I get a bunch of warnings like  
`warning: failed parsing Float64 on row=234, col=3, error=INVALID: SENTINEL, DELIMITED, INVALID_DELIMITER` Not sure what that means.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 13, 2018, 5:01am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/4 "2018-12-13T05:01:39Z")

</div>

> [@LeoK987](#):
>
> CSV.read

can you try to use CSVFiles.jl and FileIO.jl instead? Perhaps they may be more robust?

You can also try `missingstrings = "NA"` in the CSV.read?

---

<div class="post-metadata">

**Author:** ![LeoK987](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leok987/32/4880_2.png) [@LeoK987](https://discourse.julialang.org/u/LeoK987)\
**Post date:** [December 13, 2018, 5:13am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/5 "2018-12-13T05:13:07Z")

</div>

Thank you very much again. `missingstrings = "NA"` got rid of those warnings.

Can you elaborate on why CSVFiles.jl and FileIO.jl might be more robust?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 13, 2018, 6:08am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/6 "2018-12-13T06:08:02Z")

</div>

CSVFiles is not tied to any particular backend. So i think of it as having curated the right csv reader. So more likely to be robust vs someone choosing a csv reader and is new to Julia.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [December 13, 2018, 11:31am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/7 "2018-12-13T11:31:30Z")

</div>

> [@xiaodai](#):
>
> CSVFiles is not tied to any particular backend. So i think of it as having curated the right csv reader.

I don’t think this is quite right. It’s tied to`FileIO.jl` and uses `TextParse.jl` as its backend. In my experience, both packages work very well for most situations though they have slightly different interfaces - sometimes I think one of them doesn’t do something only to find out later I was using the wrong keyword.

There are a small number (decreasing all the time) of things that one does better than the other, so it is often worth giving both a shot. I tend to use CSV for everything for the pedantic reason that I don’t like `using` two packages for one functionality.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 13, 2018, 11:40am UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/8 "2018-12-13T11:40:28Z")

</div>

> [@kevbonham](#):
>
> uses `TextParse.jl` as its backend

Whats to stop CSVfiles from switching to another backend without the user noticing? Why did it choose textparse? Cos it tested a few of them and choose it. Of course it has to choose a backend but which one is almost something the user doesnt have to know

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [December 13, 2018, 5:19pm UTC](https://discourse.julialang.org/t/dataframe-has-na-what-best-to-do/18621/9 "2018-12-13T17:19:56Z")

</div>

That’s true - perhaps I misunderstood your point.
