# String31 in dataframe

**URL:** <https://discourse.julialang.org/t/string31-in-dataframe/102032>\
**Category:** New to Julia\
**Tags:** question, dataframes\
**Created:** [July 24, 2023, 5:49pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032 "2023-07-24T17:49:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Axze-rgb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/axze-rgb/32/50986_2.png) [@Axze-rgb](https://discourse.julialang.org/u/Axze-rgb)\
**Post date:** [July 24, 2023, 5:49pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032/1 "2023-07-24T17:49:24Z")

</div>

Hello, I have a 2 columns data frame. One column is Int64, which is fine. The other should be float64 I think, because it has numbers such as: 0.1254, 0.8458 etc. However, I have strange results when plotting and I realyze that typeof() returns “string31” for that float column. I guess there is some data points that are interpreted as string. Now, the problem is that the column contains 2 millions data points. How do I find the offending value(s)? I don’t even know where to start.  
Thank you

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [July 24, 2023, 5:55pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032/2 "2023-07-24T17:55:02Z")

</div>

Check out the [types](https://csv.juliadata.org/stable/reading.html#types) section of the documentation. You probably want something like `types = [Int64, Float64]` as a keyword argument. Failures will be returned as `nothing`, I think.

---

<div class="post-metadata">

**Author:** ![Axze-rgb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/axze-rgb/32/50986_2.png) [@Axze-rgb](https://discourse.julialang.org/u/Axze-rgb)\
**Post date:** [July 24, 2023, 6:04pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032/3 "2023-07-24T18:04:50Z")

</div>

Sorry, maybe I was not clear, what I want is not to change type or get rid of them (or at least not at first), I want to find the “offending” values to look at them. They might be errors, or they might not and be significant. This is what bugs me.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [July 24, 2023, 6:05pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032/4 "2023-07-24T18:05:52Z")

</div>

Check out which ones in the resulting data frame parse to `nothing` and see what row they are in.

You can also parse as a `String`, then call `tryparse` on the column to make a new one that simulates what CSV.jl is doing, then filter all the observations that have `nothing` in the float variable.

---

<div class="post-metadata">

**Author:** ![Axze-rgb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/axze-rgb/32/50986_2.png) [@Axze-rgb](https://discourse.julialang.org/u/Axze-rgb)\
**Post date:** [July 24, 2023, 6:16pm UTC](https://discourse.julialang.org/t/string31-in-dataframe/102032/5 "2023-07-24T18:16:37Z")

</div>

thanks, here is what I did (while you were answering). I thought that probably it wouldn’t succeed in parsing (okay it was just a bet).

> df2.VAF2=parse.(Float64, df2.VAF)  
> ArgumentError: cannot parse String31(" 0.21659,0.294931") as Float64

I have TWO values separated by “,”. This makes actual sense, and means I have a decision to take upstream in my pipeline! I know what’s causing this.

Thanks 🙂
