# Problem related to CSV.read() and strings

**URL:** <https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595>\
**Category:** General Usage\
**Created:** [March 8, 2018, 1:37pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595 "2018-03-08T13:37:30Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![IljaK91](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iljak91/32/44301_2.png) [@IljaK91](https://discourse.julialang.org/u/IljaK91)\
**Post date:** [March 8, 2018, 1:37pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/1 "2018-03-08T13:37:30Z")

</div>

I had a possibly related problem in the [past](https://discourse.julialang.org/t/unicode-related-error-when-reading-a-csv/7132/3).

I want to read a large data set through CSV.read(). The data set contains comments from on online forum that I scraped using Python and saved it as JSON. In Python I converted it to a unicode-encoded csv and try to read it now using CSV.read(). You can find the file [here](https://www.dropbox.com/s/q77ja7euq420ih5/output_speculation_unicode.csv?dl=0https://www.dropbox.com/s/q77ja7euq420ih5/output_speculation_unicode.csv?dl=0) (beware, it has 700 MB). This is the code that I am using:

```julia
# Define some inline functions that allow us to read comments properly
escape_double_quote(s::String) = replace(s, "\"\"", "\\\"")
escape_back_quote(s::String) = replace(s, "\\\"", "\\\\\"");
esca(s::String) = escape_double_quote(escape_back_quote(s));

# Read data from the Speculation subforum, this does not work at the moment
f = open("output_speculation_unicode.csv")
cleaned_file = IOBuffer(readstring(f))
df_speculation_raw = CSV.read(cleaned_file, DataFrame, rows_for_type_detect = 200000)

```

The first part relates to problems that were mentioned in [this thread](https://discourse.julialang.org/t/unicode-related-error-when-reading-a-csv/7132/3).

However, when I run this code, this is the preview that I receive:

![grafik](https://global.discourse-cdn.com/julialang/original/3X/7/8/78a4fe5069c5131ec28891e6353bbc8b52b777d7.png)

As you can see, something gets mixed up on the way and instead of seeing the authors name and the comment id, I see random dates in those columns. My guess is that something is mixed up in the comments with the quotation marks and I need to add even more cases of strings that I want to replace.

If anyone has an idea what I could do, I’d be very grateful!

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [March 8, 2018, 5:24pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/2 "2018-03-08T17:24:06Z")

</div>

```julia
 using CSVFiles, FileIO, DataFrames

df = load("output_speculation_unicode.csv") |> DataFrame

```

just seems to work for me on that file, without any hacks.

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [March 8, 2018, 5:41pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/3 "2018-03-08T17:41:05Z")

</div>

In light of @davidanthoff’s comment, try to make sure that you are using up-to-date versions of both DataFrames and CSV. Annoyingly, the current package manager sometimes makes that rather difficult.

---

<div class="post-metadata">

**Author:** ![IljaK91](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/iljak91/32/44301_2.png) [@IljaK91](https://discourse.julialang.org/u/IljaK91)\
**Post date:** [March 9, 2018, 11:45am UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/4 "2018-03-09T11:45:17Z")

</div>

That works, thanks a lot!

@ExpandingMan: I was using `CSV.jl` in version 0.2.2 and `DataFrames.jl` in version 0.11.5, so I think I am up-to-date.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 9, 2018, 9:41pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/5 "2018-03-09T21:41:56Z")

</div>

Interesting. Would you be able to identify a subset of rows which reproduces the problem with `CSV.read`? That would be helpful to find a fix. I guess it could be related to quoting issues.

---

<div class="post-metadata">

**Author:** ![Liso](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@Liso](https://discourse.julialang.org/u/Liso)\
**Post date:** [March 10, 2018, 12:43am UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/6 "2018-03-10T00:43:50Z")

</div>

I suspect that there is still problem with unwanted escaping (as I described [here](https://discourse.julialang.org/t/unicode-related-error-when-reading-a-csv/7132/19)):

```julia
julia> CSV.readsplitline(IOBuffer("104652,Thanks \\,a"))
2-element Array{CSV.RawField,1}:
 CSV.RawField("104652", false)      
 CSV.RawField("Thanks \\,a", false)

```

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 10, 2018, 10:22am UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/7 "2018-03-10T10:22:14Z")

</div>

@Liso AFAICT this behavior is correct. What is confusing is that `\\` actually represents a single backquote in the string, it’s doubled because that’s how you can type it in Julia:

```julia
julia> "\\"[1]
'\\': ASCII/Unicode U+005c (category Po: Punctuation, other)

julia> print("\\")
\

```

---

<div class="post-metadata">

**Author:** ![Liso](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@Liso](https://discourse.julialang.org/u/Liso)\
**Post date:** [March 10, 2018, 10:15pm UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/8 "2018-03-10T22:15:03Z")

</div>

> [@nalimilan](#):
>
> AFAICT this behavior is correct.

Please look at [Super CSV – What is CSV?](http://super-csv.github.io/super-csv/csv_specification.html)

`6. Fields containing line breaks (CRLF), double quotes, and commas should be enclosed in double-quotes.`

or

` Please note that the sometimes-used convention of escaping double-quotes as \" (instead of "") is not supported.`

Although there is possibility to implement different dialects too (for inspiration look at [csv — CSV File Reading and Writing — Python 3.12.0 documentation](https://docs.python.org/3/library/csv.html#csv-fmt-params) ), de facto standard rfc-4180 define different behavior than CSV.jl implemements!

CSVFiles, FileIO are more standard in this case (although work with dataformats or “magic” bytes is really heavy):

```julia
julia> open("/tmp/tst.csv", "w") do f write(f, IOBuffer("a,b,c\r\n104652,\"Thanks \\\",a\r\n")) end

julia> load("/tmp/tst.csv") |> DataFrame
1×3 DataFrames.DataFrame
│ Row │ a │ b │ c │
├─────┼────────┼─────────────┼─────┤
│ 1 │ 104652 │ "Thanks \\" │ "a" │

```

)

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 11, 2018, 11:20am UTC](https://discourse.julialang.org/t/problem-related-to-csv-read-and-strings/9595/9 "2018-03-11T11:20:47Z")

</div>

Ah, OK, that’s not strictly speaking a bug then, just a debatable default. See [this issue](https://github.com/JuliaData/CSV.jl/issues/166).
