# Find string with special characters in data frame read with the CSV package

**URL:** <https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967>\
**Category:** Data\
**Tags:** strings, dataframes, csv\
**Created:** [August 18, 2023, 9:47pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967 "2023-08-18T21:47:13Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![NunonuN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nunonun/32/38033_2.png) [@NunonuN](https://discourse.julialang.org/u/NunonuN)\
**Post date:** [August 18, 2023, 9:47pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/1 "2023-08-18T21:47:13Z")

</div>

Supposing I have the data frame:

```julia
using DataFrames

df = DataFrame(year = [2012, 1993, 1991, 1984, 1957, 1972, 1980], lang = ["Julia", "R", "Python", "Matlab", "Fortran", "C", "C++"])

```

I can find the creation year of C++ by executing the following command:

```julia
df[df[:, 2] .== "C++", :][1, 1]

```

However, if I read this data frame from a text file using CSV

```julia
using CSV

dfr = CSV.read("data.csv", DataFrame)

```

and perform the same search, _i.e._,

```julia
dfr[dfr[:, 2] .== "C++", :][1, 1]

```

I get a “BoundsError”, because the returned data frame is empty.  
That is, the searching command works if I build the data frame by hand, but fails if I read it from a text file using CSV.  
I also tried to use the `filter` function, but the results are similar.  
What am I doing wrong?

I’m using Julia version 1.9.1, CSV v0.10.11, and DataFrames v1.6.1.  
I’m a Manjaro user.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 18, 2023, 10:20pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/2 "2023-08-18T22:20:18Z")

</div>

I cannot reproduce your problem. I did `CSV.write("data.csv", df)` and read it back and all worked without an issue.

Also, if you were interested, this is how I would writhe the operation you perform:

```julia
julia> only(dfr.year[dfr.lang .== "C++"])
1980

```

---

<div class="post-metadata">

**Author:** ![NunonuN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nunonun/32/38033_2.png) [@NunonuN](https://discourse.julialang.org/u/NunonuN)\
**Post date:** [August 18, 2023, 10:25pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/3 "2023-08-18T22:25:57Z")

</div>

Thank you for your reply.  
In that case, I guess something is wrong with my system. I’ll try to figure that out.

And thanks for your code suggestion. It’s elegant, and I’ll be using it from now on. 👍

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [August 19, 2023, 7:21am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/4 "2023-08-19T07:21:21Z")

</div>

Can you show what `dfr` looks like after you read it in?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [August 19, 2023, 1:37pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/5 "2023-08-19T13:37:31Z")

</div>

Can you share the CSV file?

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 19, 2023, 2:34pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/6 "2023-08-19T14:34:09Z")

</div>

While waiting for the OP to clarify how the facts went, I tried some conjectures and this seems to be one probable (or possible)

```julia
str="""
2012 Julia   
1993 R
1991 Python
1984 Matlab
1957 Fortran
1972 C
1980 C++
"""

open("datacpp.csv", "w") do file
   write(file, str)
end

julia> dfr = CSV.read("datacpp.csv", DataFrame)
6×1 DataFrame
 Row │ 2012 Julia    
     │ String15       
─────┼────────────────
   1 │ 1993 R
   2 │ 1991 Python
   3 │ 1984 Matlab
   4 │ 1957 Fortran
   5 │ 1972 C
   6 │ 1980 C++

julia> dfr[dfr[:, 2] .== "C++", :][1, 1]
ERROR: BoundsError: attempt to access data frame with 1 column at index [2]

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 19, 2023, 4:10pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/8 "2023-08-19T16:10:03Z")

</div>

I know. I was just trying to reconstruct the crime scene 😄

---

<div class="post-metadata">

**Author:** ![NunonuN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nunonun/32/38033_2.png) [@NunonuN](https://discourse.julialang.org/u/NunonuN)\
**Post date:** [August 21, 2023, 4:50am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/9 "2023-08-21T04:50:07Z")

</div>

**Thank you all** , I appreciate your interest in my question.

You can download the original `.csv` file from [here](https://github.com/JuliaAcademy/DataScience/blob/main/programming_languages.csv).  
In case of need, how would you suggest I share a file with you? I guess it’s not possible to share files here in the forum, right? Any platform you recommend for this kind of cases?

My guess is that the problem has to do with the encoding of some special characters, in particular the “+” sign.  
I updated my OS today and currently only the data frame read from the original `.csv` file shows the problem. Here are the tests I ran.  
I’ll use the following helper function, to make testing easier:

```julia
function year_created(df, lang::String)
  res = df[:, 1][lowercase.(df[:, 2]) .== lowercase(lang)]
  !isempty(res) && return only(res)
  error("Could not find the programming language.")
end

```

1. Testing on the `.csv` file mentioned above (FAIL)

**NOTE** : I saved the file as “plangs01.csv”.

```julia
julia> df1 = CSV.read("plangs01.csv", DataFrame)
73×2 DataFrame
 Row │ year plang
     │ Int64 String31
─────┼───────────────────────────────────
   1 │ 1951 Regional Assembly Language
   2 │ 1952 Autocode
   3 │ 1954 IPL
  ⋮ │ ⋮ ⋮
  70 │ 2011 Red
  71 │ 2011 Elixir
  72 │ 2012 Julia
  73 │ 2014 Swift

julia> year_created(df1, "julia")
2012

julia> year_created(df1, "c++")
ERROR: Could not find the programming language.
Stacktrace:
 [1] year_created(df::DataFrame, lang::String)
   @ Main ./REPL[101]:4
 [2] top-level scope
   @ REPL[109]:1

julia> year_created(df1, "c#")
2001

```

1. Testing on a fresh data frame (PASS)

```julia
julia> df2 = DataFrame(
       year = [1993, 1991, 1984, 1957, 1972, 1980, 2012],
       lang = ["R", "Python", "MATLAB", "FORTRAN", "C", "C++", "Julia"]
       )
7×2 DataFrame
 Row │ year lang
     │ Int64 String
─────┼────────────────
   1 │ 1993 R
   2 │ 1991 Python
   3 │ 1984 MATLAB
   4 │ 1957 FORTRAN
   5 │ 1972 C
   6 │ 1980 C++
   7 │ 2012 Julia

julia> year_created(df2, "JUlia")
2012

julia> year_created(df2, "c++")
1980

```

1. Testing on a newly created `.csv` file (PASS)

```julia
julia> str = """
       "year","lang"
       1993,R
       1991,Python
       1984,Matlab
       1957,Fortran
       1972,C
       1980,C++
       2001,C#
       2012,Julia
       """
"\"year\",\"lang\"\n1993,R\n1991,Python\n1984,Matlab\n1957,Fortran\n1972,C\n1980,C++\n2001,C#\n2012,Julia\n"

julia> open("plangs03.csv", "w") do file
       write(file, str)
       end
93

julia> df3 = CSV.read("plangs03.csv", DataFrame)
8×2 DataFrame
 Row │ year lang
     │ Int64 String7
─────┼────────────────
   1 │ 1993 R
   2 │ 1991 Python
   3 │ 1984 Matlab
   4 │ 1957 Fortran
   5 │ 1972 C
   6 │ 1980 C++
   7 │ 2001 C#
   8 │ 2012 Julia

julia> year_created(df3, "julia")
2012

julia> year_created(df3, "c++")
1980

julia> year_created(df3, "c#")
2001

```

* * *

I’m curious to know what your Test 1 results are… 🤔

@rocco_sprmnt21, as you can see from my **Test 3** , your example is currently working for me, but I used to have the same problem you’re reporting.

---

<div class="post-metadata">

**Author:** ![NunonuN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nunonun/32/38033_2.png) [@NunonuN](https://discourse.julialang.org/u/NunonuN)\
**Post date:** [August 21, 2023, 5:04am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/10 "2023-08-21T05:04:24Z")

</div>

In fact, I **found the culprit**!  
By examining the `.csv` file, I realised that some lines were ending in a white space, which caused the corresponding languages to not match the input in the `year_created` function in my last post. By deleting those spaces, everything works fine. 🕶

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [August 21, 2023, 5:21am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/11 "2023-08-21T05:21:17Z")

</div>

If you are working with fumes where data was entered manually you’ll often find trailing whitespace, in this case you can use `strip` to remove it.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 21, 2023, 7:40am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/12 "2023-08-21T07:40:32Z")

</div>

> [@NunonuN](#):
>
> @rocco_sprmnt21, as you can see from my **Test 3** , your example is currently working for me, but I used to have the same problem you’re reporting.

My example was different from yours, deliberately not having used the comma (default separator) as field separator, so that the search for the string “c++”, with the criterion of exact equality, fails(\*).  
In my case, it fails because the created dataframe has only one column containing, for each row, the string “year lang”.  
It might be useful in analyzing situations of this type to loosen the matching criterion, using for example (as implicitly suggested by @Dan) the contains(str, substr) function or similar.

(\*) I tried to simulate a cut and paste operation, hypothesizing what could have happened.  
I selected the REPL output of the dataframe and pasted it embedding it in a string, to be able to save it as a text file.

---

<div class="post-metadata">

**Author:** ![NunonuN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nunonun/32/38033_2.png) [@NunonuN](https://discourse.julialang.org/u/NunonuN)\
**Post date:** [August 21, 2023, 5:31pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/13 "2023-08-21T17:31:19Z")

</div>

> [@rocco\_sprmnt21](#):
>
> My example was different from yours, deliberately not having used the comma (default separator) as field separator, so that the search for the string “c++”, with the criterion of exact equality, fails(\*).

This is interesting, because I have partially tested your code, and it had worked. I created the same string as in Test 3 above, but replaced the commas with spaces. If you try it, you’ll see that **it works as long as there is no extra space after the first line**.

To make it clearer (the difference between `stra` and `strb` is a trailing space after “lang”):

- Test A (PASS)

- Test B (FAIL)

I’m not a software engineer, but I wonder if this behaviour is expected from a programming language or if Julia should be made “more robust” relatively to this kind of small differences that might be hard to catch…

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [August 21, 2023, 6:31pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/14 "2023-08-21T18:31:23Z")

</div>

I believe that the CSV package (like others) uses heuristics to be as convenient as possible.  
For example in the case in question if it finds a list of lines with the same “structure” [Number, Spaces, Word] it will think it is doing what it likes by interpreting the text as two columns of data separated by spaces.

If it finds any of these lines that have some extra trailing spaces, it can’t arbitrarily split into columns and put everything together, as was the case in my case as well [looking closer I saw that there were trailing spaces in one of the lines].

Reading the CSV documentation I believe these rules should be made explicit. Otherwise, you can ask the package maintainers about it.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [August 21, 2023, 10:35pm UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/15 "2023-08-21T22:35:54Z")

</div>

> [@NunonuN](#):
>
> Test B (FAIL)

You could read using space delimiter: `CSV.read(file, DataFrame, delim=" ")` .

This will create a column of missings that can be cleaned out:

```julia
using CSV, DataFrames
dfb = CSV.read("testb.csv", DataFrame, delim=" ")
dfb[!, Not(all.(ismissing, eachcol(dfb)))]

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [August 22, 2023, 1:03am UTC](https://discourse.julialang.org/t/find-string-with-special-characters-in-data-frame-read-with-the-csv-package/102967/16 "2023-08-22T01:03:43Z")

</div>

An actual use of LLMs (using Llama2-13B):

```julia
convert the following into a valid CSV table:
"year" "lang"
1993 R
1991 Python
1980 C++
2012 Julia

```

Output:

```julia
Sure! Here's the valid CSV table:

year,lang
1993,R
1991,Python
1980,C++
2012,Julia

```
