# Regex, PCRE2, and the \`PCRE2\_UCP\` / \`(\*UCP)\` flag

**URL:** <https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930>\
**Category:** Internals & Design\
**Tags:** strings, regex\
**Created:** [May 16, 2018, 3:15pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930 "2018-05-16T15:15:30Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 16, 2018, 3:15pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/1 "2018-05-16T15:15:30Z")

</div>

In the discussion on GitHub: [#27084](https://github.com/JuliaLang/julia/issues/27084#issuecomment-389506415) it seems people are gravitating towards setting the `PCRE2_UCP` flag by default.

I would strongly recommend against that, and from their comments, it looks like the authors of PCRE2 also seem to feel the same way, see [PCRE2 documentation](http://www.pcre.org/current/doc/html/pcre2pattern.html).

> By default, characters whose code points are greater than 127 never match \d, \s, or \w, and always match \D, \S, and \W, … These escape sequences retain their original meanings from before Unicode support was available, mainly for efficiency reasons.

> Matching these sequences is noticeably slower when PCRE2\_UCP is set.

> Matching characters by Unicode property is not fast, because PCRE2 has to do a multistage table lookup in order to find a character’s property. That is why the traditional escape sequences such as \d and \w do not use Unicode properties in PCRE2 by default, though you can make them do so by setting the PCRE2\_UCP option or by starting the pattern with (\*UCP).

Note that if you really want to take the performance hit, and use the Unicode tables, then you can simply add the  
`(*UCP)` at the beginning of the pattern, or if you only want to use the Unicode tables in certain parts of the regex, you can use the following equivalences:

> \d any character that matches \p{Nd} (decimal digit)  
> \s any character that matches \p{Z} or \h or \v  
> \w any character that matches \p{L} or \p{N}, plus underscore

i.e. instead of `\d`, explicitly write `\{Nd}`, instead of `\s`, `[\pZ\h\v]`, and instead of `\w`, `[\pL\pN_]`.  
If that is deemed too inconvenient, we could simply preprocess the pattern, and make replacements such as:  
(or whatever escape sequence would be easy to remember and use)

- `\ñ` =\> `\P{Nd}` (n for numeric, tilde to indicate Unicode)
- `\Ñ` =\> `\P{Nd}`
- `\á` =\> `[\pZ\h\v]` (whitespace)
- `\Á` =\> `[^\pZ\h\v]`
- `\í` =\> `[\pL\pN_]` (i for identifier, accent for Unicode)
- `\Í` =\> `[^\pL\pN_]`

The problem with making it the default, is that there is no easy way to _disable_ it, if you want the best performance, and you are only concerned with matching ASCII characters anyway).

I made this more convenient for by adding a `u` option (in addition to the `i`, `m`, `s`, `x` ones already available), for [StrRegex.jl](https://github.com/JuliaString/StrRegex.jl).

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [May 16, 2018, 3:37pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/2 "2018-05-16T15:37:18Z")

</div>

> [@ScottPJones](#):
>
> The problem with making it the default, is that there is no easy way to _disable_ it, if you want the best performance, and you are only concerned with matching ASCII characters anyway).

AFAICT we just need to pass `PCRE2_NEVER_UCP`, which can be exposed to the user via a flag.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 16, 2018, 4:10pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/3 "2018-05-16T16:10:03Z")

</div>

IIUC, that just disallows setting UCP in the regex pattern via `(*UCP)`, and doesn’t disable UCP if it’s been set in the options.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [May 17, 2018, 9:00am UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/4 "2018-05-17T09:00:41Z")

</div>

Actually we can just avoid passing `PCRE2_UCP` when we get the flag, so I don’t see the problem.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 11:53am UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/5 "2018-05-17T11:53:12Z")

</div>

I think it boils down to:

1. Should the much faster case be the default, given that much of the usage of regex strings will be in cases where you don’t need the Unicode tables
2. Should the use of the tables and associated performance degradation be opt-in (as PCRE2 does) (by simply adding a ‘u’ to the regex options, as I’ve already done for StrRegex), or opt-out (as you have suggested)?
3. Does Julia want to have it’s PCRE2 binding differ from the defaults of the well-documented and heavily used PCRE2 library itself?

In the last few years, has anybody before complained about Julia’s _not_ using the Unicode tables by default for  
`\w`, `\s`, and `\d` ?

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 12:20pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/6 "2018-05-17T12:20:26Z")

</div>

> [@nalimilan](#):
>
> Actually we can just avoid passing `PCRE2_UCP` when we get the flag, so I don’t see the problem.

Also, a major point would be that changing the default for Julia regexes would be a breaking change,  
one that could not easily be handled via a deprecation warning.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [May 17, 2018, 1:15pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/7 "2018-05-17T13:15:30Z")

</div>

I’d rather be Unicode-aware (and therefore more correct in most situations) by default, and faster with a flag. This is similar to what we do with e.g. `@inbounds` or `@simd`. The survey of other languages made in the issue has shown that there is no standard behavior across languages and regex libraries.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 17, 2018, 1:54pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/8 "2018-05-17T13:54:24Z")

</div>

> [@nalimilan](#):
>
> The survey of other languages made in the issue has shown that there is no standard behavior across languages and regex libraries.

Then, if it’s true that it would be a breaking change, don’t do it.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 4:01pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/9 "2018-05-17T16:01:54Z")

</div>

> [@nalimilan](#):
>
> The survey of other languages made in the issue has shown that there is no standard behavior across languages and regex libraries.

Python, Rust, Go all have regex support in separate packages/crates/classes, _not_ baked into the base language (which is what I believe should done for Julia before v1.0 is released).

For those, the defaults can be different depending on which library is being used for the regex support.

Programs using the PCRE2 library directly, expect the default to be ASCII for those 4 characters (and allow setting the PCRE2\_UCP flag, or having `(*UCP)` at the beginning of a pattern).  
Programs using the ICU library directly, word boundaries are ASCII by default, a flag can be set to allow slower/more complex Unicode aware word boundaries.  
Swift uses the NSRegularExpression class (which uses ICU internally), word boundaries by default ASCII.  
Java uses is like PCRE2 (and Julia up til now), ASCII by default, but has a regex compilation flag to make those 4 (and POSIX character classes) Unicode aware.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [May 17, 2018, 4:13pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/10 "2018-05-17T16:13:52Z")

</div>

> [@ScottPJones](#):
>
> Programs using the ICU library directly, word boundaries are ASCII by default, a flag can be set to allow slower/more complex Unicode aware word boundaries.  
> Swift uses the NSRegularExpression class (which uses ICU internally), word boundaries by default ASCII.

[AFAICT](http://userguide.icu-project.org/strings/regexp) in ICU `\w` matches any Unicode word character, doesn’t it?

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 4:33pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/11 "2018-05-17T16:33:32Z")

</div>

> [@nalimilan](#):
>
> [AFAICT](http://userguide.icu-project.org/strings/regexp) in ICU `\w` matches any Unicode word character, doesn’t it?

Ah, my mistake!  
I was thinking of the `w` _option_, in ICU, i.e.

> Controls the behavior of \b in a pattern. If set, word boundaries are found according to the definitions of word found in Unicode UAX 29, Text Boundaries. By default, word boundaries are identified by means of a simple classification of characters as either “word” or “non-word”, which approximates traditional regular expression behavior. The results obtained with the two options can be quite different in runs of spaces and other non-word characters.`

I should have gone back and RTFM to refresh my memory on that case! 😂

---

<div class="post-metadata">

**Author:** ![jeff.bezanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff.bezanson/32/48_2.png) [@jeff.bezanson](https://discourse.julialang.org/u/jeff.bezanson)\
**Post date:** [May 17, 2018, 6:56pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/12 "2018-05-17T18:56:46Z")

</div>

I don’t know; it seems pretty weird to consider non-ASCII characters non-{word, digit} by default. Why not support all human languages in Unicode by default? Code suddenly failing e.g. when one day a filename is not in English doesn’t seem like a good default. We can certainly add a flag to enable ASCII mode if desired though.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 7:11pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/13 "2018-05-17T19:11:21Z")

</div>

Well, as an example, non-ASCII digits are problematic, usually people (even in other languages) don’t really want anything but 0-9, as well as taking the performance hit by default.  
It might be good to have an option to toggle that separately.

Maybe have both an `a` and a `u` compile option flags, along with defaults based on the string type, would be OK.  
So, `ASCIIStr`, `Text1Str`, `Text2Str`, `Text4Str` would not by default enable `PCRE2_UCP`, but could enable with `u`), and `String`, `LatinStr`, `UTF8Str`, `UCS2Str` etc. would (but could disable with `a`).

Does that seem like a reasonable approach?

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 7:13pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/14 "2018-05-17T19:13:30Z")

</div>

> [@jeff.bezanson](#):
>
> Code suddenly failing e.g. when one day a filename is not in English doesn’t seem like a good default.

True, although nobody has complained yet about that happening in Julia, and the regex’s in base use `\w`, it should be fixed.

---

<div class="post-metadata">

**Author:** ![jeff.bezanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff.bezanson/32/48_2.png) [@jeff.bezanson](https://discourse.julialang.org/u/jeff.bezanson)\
**Post date:** [May 17, 2018, 7:19pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/15 "2018-05-17T19:19:21Z")

</div>

Yes that seems reasonable. But is it possible/efficient to specify the option per string to match, and not just at regex compile time?

---

<div class="post-metadata">

**Author:** ![jeff.bezanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff.bezanson/32/48_2.png) [@jeff.bezanson](https://discourse.julialang.org/u/jeff.bezanson)\
**Post date:** [May 17, 2018, 7:20pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/16 "2018-05-17T19:20:13Z")

</div>

> [@ScottPJones](#):
>
> and the regex’s in base use `\w` , it should be fixed.

True, but this is also a case in point — people expect `\w` to match word characters, not just English word characters!

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [May 17, 2018, 7:24pm UTC](https://discourse.julialang.org/t/regex-pcre2-and-the-pcre2-ucp-ucp-flag/10930/17 "2018-05-17T19:24:28Z")

</div>

> [@jeff.bezanson](#):
>
> Yes that seems reasonable. But is it possible/efficient to specify the option per string to match, and not just at regex compile time?

It’s not something that really depends on the string being matched (except that that can determine the default value to use)

You have to do tricks like I did in StrRegex, otherwise you’d be constantly recompiling.  
It can be done though, it would need to compile the pattern both ways (probably lazily for the non-default cases).

I’ll make sure this is handled efficiently in StrRegex.

Basically, if you have an `UTF8Str` or `UniStr`, the default would be to have `u` set for the pattern, but not for `ASCIIStr` (similar to the `PCRE2_UTF`) flag, and at run-time, you pick the compiled regex to use based on the the type of the string being matched and the options given for the pattern, and if that field is `C_NULL`, you compile the regex, with those options, and the correct code unit size, and then call the appropriate match code.
