# Regex replace by group index

**URL:** https://discourse.julialang.org/t/regex-replace-by-group-index/131342
**Category:** General Usage
**Tags:** regex
**Created:** [August 4, 2025, 4:57pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342 "2025-08-04T16:57:41Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![paiv](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paiv/32/218059_2.png) [@paiv](https://discourse.julialang.org/u/paiv)
#### Post date: [August 4, 2025, 4:57pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/1 "2025-08-04T16:57:41Z")

</div>

When doing string replace with regex, I want to know which capture group matched, so I can perform replacement specific to that group.

For example, if matched at word boundary, the letter should be replaced with `X`, and if followed by `c`, the letter should be replaced with `_`:

```julia
# "abac" => "Xb_c"
tr = ["X", "_"]
replace("abac", r"\b(\w)|(\w)(?=c)" => f)
f(s) = tr[1] # ?

```

Here `f` would be given the string `"a"` in both matches, and I have no indication of which group matched. Splitting replace into multiple runs is not possible, since the context for each match will change.

Currently I am working around this by providing custom method

```julia
Base._replace(io::IO, repl_s::_MyT, str, r, re::Base.RegexAndMatchData) = begin
    n = Base.PCRE.substring_length_bynumber(re.match_data, 1)
    ...

```

but I don’t feel at ease by patching internal methods.

Is there a better way? For reference, other languages would provide full `Match` object to `f` in this case.

---

<div class="post-metadata">

### Author: ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)
#### Post date: [August 4, 2025, 6:44pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/2 "2025-08-04T18:44:54Z")

</div>

Welcome to Julia! I’m not the best person to answer, and I don’t think I have your answer here, but I’ll do what I can to start.

It doesn’t seem to be quite what you’re after, but the [`SubstitutionString`](https://docs.julialang.org/en/v1/base/strings/#Base.SubstitutionString) possibility in [`replace`](https://docs.julialang.org/en/v1/base/strings/#Base.replace-Tuple%7BIO,%2520AbstractString,%2520Vararg%7BPair%7D%7D) looks like it could _maybe_ relevant.

I don’t know if it’s applicable to your actual problem, but in this simple example you can just apply multiple replacement patterns:

```julia-repl
julia> replace("abac", r"\b(\w)" => "X", r"(\w)(?=c)" => "_")
"Xb_c"

```

There might also be something you could do with `eachmatch`, but it might be some work to plumb that into a `replace`-like function to actually make the substitution based on what it matched.

```julia-repl
# The RegexMatches contain more info than they show here, although some is likely internal.
# Try calling `dump` on one.
julia> eachmatch(r"\b(\w)|(\w)(?=c)", "abac") |> collect
2-element Vector{RegexMatch}:
 RegexMatch("a", 1="a", 2=nothing)
 RegexMatch("a", 1=nothing, 2="a")

```

There might be room to add a feature to `replace` where you could make the function be passed the entire `RegexMatch` object rather than only the `SubString`. That seems like it is maybe what it should have been to begin with.

---

<div class="post-metadata">

### Author: ![eldee](https://avatars.discourse-cdn.com/v4/letter/e/b5a626/32.png) [@eldee](https://discourse.julialang.org/u/eldee)
#### Post date: [August 4, 2025, 7:32pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/3 "2025-08-04T19:32:04Z")

</div>

I also suspect there is a way to use `replace`, but alternatively, you could use `findall`. It seems a bit trickier than I had originally imagined though:

```julia-repl
julia> function replace_regex(str, regex_replacements)
           # e.g. for str = "abbc", regex_replacements = [r"(\w)(?=c)" => "_", r"\b(\w)" => "X"]
           ranges_replacements = map(regex_replacements) do (rgx, repl)
               return findall(rgx, str) .=> repl
           end # Vector{Vector{Pair}}, e.g. [[2:2 => "_", 3:3 => "_"], [1:1 => "X"]]
           ranges_replacements = sort(reduce(vcat, ranges_replacements)) # Vector{Pair}, e.g. [1:1 => "X", 2:2 => "_", 3:3 => "_"]
           # Note: if the ranges overlap, things will go wrong. 
           #
           new_str = str
           for (range, repl) in Iterators.reverse(ranges_replacements)
               # Iterate from back to front, so that the ranges remain valid.
               new_str = new_str[begin:range.start-1] * repl * new_str[range.stop+1:end]
           end
           return new_str
       end;

julia> replace_regex("abac", [r"\b(\w)" => "X", r"(\w)(?=c)" => "_"])
"Xb_c"

julia> replace_regex("abcc", [r"(\w)(?=c)" => "_", r"\b(\w)" => "X"])
"X__c"

julia> replace_regex("abzc", [r"z" => "c", r"(\w)(?=c)" => "_"]) # Different from replace_regex("abcc", [r"(\w)(?=c)" => "_"]): "a__c"
"ab_c"

```

---

<div class="post-metadata">

### Author: ![paiv](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paiv/32/218059_2.png) [@paiv](https://discourse.julialang.org/u/paiv)
#### Post date: [August 4, 2025, 7:33pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/4 "2025-08-04T19:33:52Z")

</div>

To augment original example, given the same regex `r"\b(\w)|(\w)(?=c)"`, the replacement should be

```julia-auto
"abac" => "Xb_c"
"ebec" => "Yb:c"
"oboc" => "Zb.c"

```

(actually dozens more)

Iterating over `eachmatch` seems to be equivalent to what I’m doing if I change regex a little to `r"\b(\w)|(\w)(?=c)|(.)"`, I’ll give it a try. Thank you.

---

<div class="post-metadata">

### Author: ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)
#### Post date: [August 4, 2025, 8:01pm UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/5 "2025-08-04T20:01:18Z")

</div>

If all you needed was a simple library of replacements you can consider something like

```julia-repl
julia> wordboundarydict = Dict("a" => "X", "e" => "Y", "o" => "Z"); # replacements after word boundary

julia> beforecdict = Dict("a" => "_", "e" => ":", "o" => "."); # replacements before c

julia> substitute(dict) = key -> get(dict, key, key); # access a key if possible, else returns the key

julia> replace("abac", r"\b(\w)" => substitute(wordboundarydict), r"(\w)(?=c)" => substitute(beforecdict))
"Xb_c"

julia> replace("ebec", r"\b(\w)" => substitute(wordboundarydict), r"(\w)(?=c)" => substitute(beforecdict))
"Yb:c"

julia> replace("oboc", r"\b(\w)" => substitute(wordboundarydict), r"(\w)(?=c)" => substitute(beforecdict))
"Zb.c"

julia> replace("xbec", r"\b(\w)" => substitute(wordboundarydict), r"(\w)(?=c)" => substitute(beforecdict)) # no replacement for x
"xb:c"

```

But otherwise it sounds like you have a plan you can use based on `eachmatch`. Good luck!

_EDIT: I finally took a look into the `Base._replace` function you were modifying. Comments around that file suggest that packages might extend some of those functions. This suggests they probably won’t change a lot, although I don’t see anything quite suggesting a guarantee. All-in-all, it looks like your approach there (with your custom types to avoid piracy) was mostly okay if that ends up being the nicest way._

---

<div class="post-metadata">

### Author: ![paiv](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paiv/32/218059_2.png) [@paiv](https://discourse.julialang.org/u/paiv)
#### Post date: [August 5, 2025, 8:11am UTC](https://discourse.julialang.org/t/regex-replace-by-group-index/131342/6 "2025-08-05T08:11:59Z")

</div>

Splitting original regex into parts and using multiple patterns interface works for me fine.
