# Non greedy regex match

**URL:** <https://discourse.julialang.org/t/non-greedy-regex-match/101268>\
**Category:** General Usage\
**Tags:** regex\
**Created:** [July 6, 2023, 4:08pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268 "2023-07-06T16:08:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![filchristou](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/filchristou/32/26760_2.png) [@filchristou](https://discourse.julialang.org/u/filchristou)\
**Post date:** [July 6, 2023, 4:08pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268/1 "2023-07-06T16:08:18Z")

</div>

How can I match non-greedily a regex in julia ?

A [general stackoverflow post](https://stackoverflow.com/questions/11898998/how-can-i-write-a-regex-which-matches-non-greedy) mentions `.*?` but it doesn’t work.

Julia docs point to [pcre2 site](https://www.pcre.org/current/doc/html/pcre2syntax.html) that mentions ` (?U) default ungreedy (lazy)` but I am not sure how to use it.

### Example

```julia-auto
julia> str = """<a href="files/Zamren.gml">GML</a> <a href="files/Zamren.graphml">GraphML</a>""";

julia> reg = r"<a href=\"files/(.*?)\.graphml\">";

julia> match(reg, str)
RegexMatch("<a href=\"files/Zamren.gml\">GML</a> <a href=\"files/Zamren.graphml\">", 1="Zamren.gml\">GML</a> <a href=\"files/Zamren")

```

The matching result should be `RegexMatch("<a href=\"files/Zamren.graphml\">", 1="Zamren")`

#### PS

Please don’t suggest custom workarounds like:

```julia-auto
julia> match(r"<a href=\"files/([^>]*)\.graphml\">", str)
RegexMatch("<a href=\"files/Zamren.graphml\">", 1="Zamren")

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 6, 2023, 4:31pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268/2 "2023-07-06T16:31:51Z")

</div>

You shouldn’t be using RegEx to parse HTML

> <https://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags/1732454#1732454>

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [July 6, 2023, 4:36pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268/3 "2023-07-06T16:36:16Z")

</div>

True, but the example should still work, no?

Edit: actually, no, the behavior is perfectly sensible. The `.*?` pattern is non-greedy, but there’s really no way for the regex-engine to know that you don’t want the match to start as early as possible. Just because `.*?` occurs somewhere in the regex doesn’t mean “make the _entire_ match as short as possible”.

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [July 6, 2023, 4:46pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268/4 "2023-07-06T16:46:29Z")

</div>

The regex-engine is going to start from the left and find the initial `<a href="files/`. After that, it will keep adding _the minimum number of letters_ (since `.*?` is indeed non-greedy) to complete the match, which gives you a total match of `"<a href=\"files/Zamren.gml\">GML</a> <a href=\"files/Zamren.graphml\">"`. So, the syntax for non-greedy matches is indeed `*?`, but in this case, “non-greedy” doesn’t do what you think it does (find a shorter match later in the string). An easy mistake to make (took me a while, too), and why regexes are often quite tricky.

I would strongly second using a proper HTML parser instead of regexes.

---

<div class="post-metadata">

**Author:** ![filchristou](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/filchristou/32/26760_2.png) [@filchristou](https://discourse.julialang.org/u/filchristou)\
**Post date:** [July 6, 2023, 5:46pm UTC](https://discourse.julialang.org/t/non-greedy-regex-match/101268/5 "2023-07-06T17:46:18Z")

</div>

ah. you are right. `*?` works indeed as expected in more appropriate situations.

```julia
julia> match(r"<a.*>", str)
RegexMatch("<a href=\"files/Zamren.gml\">GML</a> <a href=\"files/Zamren.graphml\">GraphML</a>")

julia> match(r"<a.*?>", str)
RegexMatch("<a href=\"files/Zamren.gml\">")

```

Thanks!
