# Regex matching when string has non-ascii unicode

**URL:** <https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311>\
**Category:** General Usage\
**Tags:** strings, regex, unicode\
**Created:** [October 25, 2021, 1:29am UTC](https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311 "2021-10-25T01:29:05Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![sleak](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sleak/32/15447_2.png) [@sleak](https://discourse.julialang.org/u/sleak)\
**Post date:** [October 25, 2021, 1:29am UTC](https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311/1 "2021-10-25T01:29:05Z")

</div>

I’m using regex to split a string and getting unexpected results when the input data contains non-ascii characters (specifically, a unicode \xe8). In python I can do:

```julia
>>> import re
>>> s = "caf\xe8"
>>> re.findall('.', s)
['c', 'a', 'f', 'è']

```

But in Julia (1.7rc2) I’m not able to match the e-with-grave:

```julia
julia> s = "caf\xe8"
"caf\xe8"

julia> eachmatch(r".", s) |> collect
3-element Vector{RegexMatch}:
 RegexMatch("c")
 RegexMatch("a")
 RegexMatch("f")

julia> eachmatch(r"\X", s) |> collect
3-element Vector{RegexMatch}:
 RegexMatch("c")
 RegexMatch("a")
 RegexMatch("f")

```

Am I doing something wrong, or is this a bug in the PCRE module that I should report?  
thanks,  
Steve

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 25, 2021, 2:57am UTC](https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311/2 "2021-10-25T02:57:17Z")

</div>

> [@sleak](#):
>
> But in Julia (1.7rc2) I’m not able to match the e-with-grave:
> 
> ```julia
> julia> s = "caf\xe8"
> 
> ```

This is not e-with-grave in Unicode. You are using an obsolete Latin1 encoding, probably [Windows 1252](https://en.wikipedia.org/wiki/Windows-1252), not Unicode (`\x` is _byte_ escaping, not Unicode _character_ escaping). Native Julia strings are Unicode (though you can convert from other encodings using [StringEncodings.jl](https://github.com/JuliaStrings/StringEncodings.jl)).

In Unicode, e-with-grave is [U+00E8](https://www.fileformat.info/info/unicode/char/00e8/index.htm), which can be entered as

```julia
julia> s = "caf\ue8"
"cafè"

```

(or simply as `s = "cafè"`). Notice that it prints correctly with the accent, unlike `"caf\xe8"`. Regex then works fine:

```julia
julia> eachmatch(r".", s) |> collect
4-element Vector{RegexMatch}:
 RegexMatch("c")
 RegexMatch("a")
 RegexMatch("f")
 RegexMatch("è")

```

Beware that there are subtleties with accented characters in Unicode, because they can typically be encoded in two [canonically equivalent](https://en.wikipedia.org/wiki/Unicode_equivalence) ways: as an accented character like U+00E8, _or_ as an unaccented letter followed by a “combining character” encoding the accent. In particular, the encoding

```julia
julia> s = "cafe\u0300"
"cafè"

```

using the [U+0300](https://www.fileformat.info/info/unicode/char/0300/index.htm) combining accent, is an equivalent way (according to Unicode) to express `"cafè"`, but it contains 5 “characters” (Unicode codepoints) instead of 4. Regex still works, but then the accent gets matched separately:

```julia
julia> eachmatch(r".", s) |> collect
5-element Vector{RegexMatch}:
 RegexMatch("c")
 RegexMatch("a")
 RegexMatch("f")
 RegexMatch("e")
 RegexMatch("̀")

```

(This is a property of Unicode, not Julia.)

---

<div class="post-metadata">

**Author:** ![sleak](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sleak/32/15447_2.png) [@sleak](https://discourse.julialang.org/u/sleak)\
**Post date:** [October 25, 2021, 5:22am UTC](https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311/3 "2021-10-25T05:22:32Z")

</div>

ah, that explains it, thanks! Reading the input data through StringEncodings.jl worked, now the regex matches successfully

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 26, 2021, 1:04pm UTC](https://discourse.julialang.org/t/regex-matching-when-string-has-non-ascii-unicode/70311/4 "2021-10-26T13:04:42Z")

</div>

> [@stevengj](#):
>
> using the [U+0300](https://www.fileformat.info/info/unicode/char/0300/index.htm) combining accent, is an equivalent way (according to Unicode) to express `"cafè"` , but it contains 5 “characters” (Unicode codepoints) instead of 4. Regex still works, but then the accent gets matched separately:

I should mention that this can be handled by using `\X` instead of `.` in your regex — the `\X` escape in a regular expression matches a Unicode “grapheme”, which corresponds more closely to a human-perceived “character”. (e.g. it includes a letter followed by any number of combining modifiers.)

For example:

```julia
julia> s = "cafe\u0300"
"cafè"

julia> eachmatch(r"\X", s) |> collect
4-element Vector{RegexMatch}:
 RegexMatch("c")
 RegexMatch("a")
 RegexMatch("f")
 RegexMatch("è")

```

which contains 4 matches as expected — but the final `RegexMatch("è")` match actually consists of a string `"e\u0300"` of two characters (two Unicode codepoints).
