# Regx what wrong?

**URL:** <https://discourse.julialang.org/t/regx-what-wrong/46059>\
**Category:** Offtopic\
**Created:** [September 4, 2020, 1:36pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059 "2020-09-04T13:36:52Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 4, 2020, 1:36pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/1 "2020-09-04T13:36:52Z")

</div>

Why last row is wrong ?

```julia
v=["sp. z o.o. asdas"
"sp. z o.o asdasd" 
"sp. z oo asdas"
"sp. zo.o. asdasd"
"sp. zoo. asddfa"
"sp. z o.o. afdasf"
"sp zoo. afdasf"
"sp.zoo. afdasf"
"spzoo afdasf"]

julia> occursin.(r"sp.+z.+o", v)
9-element BitArray{1}:
 1
 1
 1
 1
 1
 1
 1
 1
 0

```

Thx Paul

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [September 4, 2020, 1:42pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/2 "2020-09-04T13:42:10Z")

</div>

> [@programista](#):
>
> `occursin.(r"sp.+z.+o", v)`

`+` is 1 or more times. Between sp and z there is 0 times anything, so it fails.  
You may use `*` for 0 or more times, like:  
`occursin.(r"sp.*z.+o", v)`

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 4, 2020, 1:53pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/3 "2020-09-04T13:53:05Z")

</div>

> [@oheil](#):
>
> occursin.(r"sp.\*z.+o", v)

Thanks!  
but now is 1 row more : “zespol szkol w przem” , the last one. This solution with las row is wrong. How to find only simillar sp. z o.o. I thing space betewen Chars must be no loenger then 1-2 place. How to do ?

```julia-auto
Thanks, stars !
But julia> v=["sp. z o.o. asdas"
       "sp. z o.o asdasd"
       "sp. z oo asdas"
       "sp. zo.o. asdasd"
       "sp. zoo. asddfa"
       "sp. z o.o. afdasf"
       "sp zoo. afdasf"
       "sp.zoo. afdasf"
       "spzoo afdasf"
       "zespol szkol w przem"]

julia> occursin.(r"sp.*z.+o", v)
10-element BitArray{1}:
 1
 1
 1
 1
 1
 1
 1
 1
 1
 1

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 4, 2020, 1:53pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/4 "2020-09-04T13:53:50Z")

</div>

Have you tested this on [Regex101.com](http://Regex101.com)? Always super helpful for debugging regex

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 4, 2020, 1:58pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/5 "2020-09-04T13:58:38Z")

</div>

thanks I still practice it at [https://regexr.com/](https://regexr.com/) but it’s not easy  
🙂  
Paul

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [September 4, 2020, 2:00pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/6 "2020-09-04T14:00:48Z")

</div>

What is exactly your desired outcome?  
E.g. `.` is part of your regex meaning “any character” and part of the strings in the array. So it’s not clear what you want to match.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [September 4, 2020, 2:52pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/7 "2020-09-04T14:52:00Z")

</div>

You mean that they must start with “sp”? In this case, just use `^` as the first character of the regex. This means the regex must match from the start of the string. “zespol szkol w przem” is currently being matched because the substring “spol szko” matches (i.e., “ze **spol szko** l w przem”).

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [September 4, 2020, 4:12pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/8 "2020-09-04T16:12:48Z")

</div>

See

> **[RegexOne - Learn Regular Expressions - Lesson 2: The Dot](https://regexone.com/lesson/wildcards_dot)**
>
> RegexOne provides a set of interactive lessons and exercises to help you learn regular expressions

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 5, 2020, 3:45pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/9 "2020-09-05T15:45:16Z")

</div>

Thanks, ^ works… but the fraze can be evrywhere…  
At the monet I gave new patern: r"sp.+z.+o.{1}" works wrong in 5 and 6 rows… Please

```julia

julia> v=["sf asd sp. z o.o. asdas "
       "dfs sp. z o.o asdasd "
       "ss sp.zoo. afdasf"
       "ssssp.zoo. afdasf"
       "ss spzoo afdasf"
       "ss ds zespol szkol w przem"]
6-element Array{String,1}:
 "sf asd sp. z o.o. asdas "
 "dfs sp. z o.o asdasd "
 "ss sp.zoo. afdasf"
 "ssssp.zoo. afdasf"
 "ss spzoo afdasf"
 "ss ds zespol szkol w przem"

julia>

julia> occursin.(r"sp.+z.+o.{1}", v)
6-element BitArray{1}:
 1
 1
 1
 1
 0
 1

```

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [September 5, 2020, 4:14pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/10 "2020-09-05T16:14:29Z")

</div>

I am not sure what is your question. Are you asking why `$` makes the regex to not match any of the strings? This happens because you have defined that the second-to-last character is an `o`, what is not true for any of the strings. Did you mean to use `r"^sp.*z.+o.*$"`? I do not think there is a reason to add an `$` if you gonna use a `.*` (or `.+`) after it.

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 5, 2020, 4:20pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/11 "2020-09-05T16:20:19Z")

</div>

In my language is very importand offical shortcut : “sp. z o. o.” bat people makes many mistakes ;)I have to find every mistake combination like:: spzoo sp. zoo …  
At the moment i found this solution

Why works wrong with 10 rows? (row 9 a can remove in second step)

```julia
julia> v=["sf asd sp. z o.o. asdas "
       " dfs sp. z o.o asdasd "
       " ds sp. .z. oo asdas"
       "dsfs sp. zo.o. asdasd"
       "d sp. zoo. asddfa"
       "s sp. z o.o. afdasf"
       "ss sp.zoo. afdasf"
       "ssssp.zoo. afdasf"
       "ss spzoo afdasf"
       "ss ds zespol szkol w przem"]

julia> occursin.(r"s?p.+z.+o.{0}",v)
10-element BitArray{1}:
 1
 1
 1
 1
 1
 1
 1
 1
 0
 1

```

Paul

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [September 5, 2020, 4:37pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/12 "2020-09-05T16:37:16Z")

</div>

[https://regex101.com](https://regex101.com) maybe try to debug with one of those online regex visualized editor

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [September 5, 2020, 4:52pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/13 "2020-09-05T16:52:29Z")

</div>

I think what you are looking for is

```julia
julia> spzoo = r"s[.\s]*p[.\s]*z[.\s]*o[.\s]*o\.?";

julia> occursin.(spzoo, v)
10-element BitArray{1}:
 1
 1
 1
 1
 1
 1
 1
 1
 1
 0

```

The `[.\s]*` part allows for optional `.` and whitespace inbetween the characters. `\.?` is not strictly necessary for `occursin` but for the match it will return the trailing `.` if its in the string.

To see what this matches exactly:

```julia
julia> [match(spzoo, s).match for s in v if !isnothing(match(spzoo, s))] 
9-element Array{SubString{String},1}:
 "sp. z o.o."
 "sp. z o.o"
 "sp. .z. oo"
 "sp. zo.o."
 "sp. zoo."
 "sp. z o.o."
 "sp.zoo."
 "sp.zoo."
 "spzoo"

```

EDIT:  
You might want to add `i` after the regex to make it case insensitive

```julia
spzoo2 = r"s[.\s]*p[.\s]*z[.\s]*o[.\s]*o\.?"i;
julia> occursin(spzoo2, "s Sp. Z o.O. afdasf")
true

```

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 5, 2020, 5:18pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/14 "2020-09-05T17:18:14Z")

</div>

big help and good lesson Thanks

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 6, 2020, 2:35pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/17 "2020-09-06T14:35:26Z")

</div>

To many line …😕 I need only first 5 lines

```julia
rx=r"\b\d{2}-\d{3}\b"

julia> baza1[occursin.(rx,baza1)]
1077386-eleme "45-367"
 "45-367"
 "a 45-367"
 "0 45-367 0" 
"a 45-367 b"
 "tel. 91-321 28 81"
 "58-531 54 52"
 "58-531 54 52"
 "58-531 54 52"
 "58-531 54 52"
 "58-531 54 52"
 "91-321 28 81"
 "12-289 13"
 "12-289 13 31"
 "12-289 13 32"
 "12-289 13 31"
 "12-289 13 31"
 "67-286 24 80"
 "12-289 13 31"
...

```

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [September 6, 2020, 2:48pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/18 "2020-09-06T14:48:50Z")

</div>

The regex below only matches if the `dd-ddd` is at most preceded by any character and a space and/or followed by a space and any character.

```julia
rx=r"^(. )?\d{2}-\d{3}( .)?$"

```

This is the pattern I have seen at least. In your initial regex you considered that the extremities could only have base 10 digits (this is what the `\d` means), but in the 5 first lines you have lines in which the extremities are letters like `a` and `b` (should this be hex?).

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [September 6, 2020, 3:07pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/19 "2020-09-06T15:07:54Z")

</div>

W dniu 2020-09-06 o 16:53, Henrique Becker via JuliaLang pisze:

hex no! All data are just string UTF8

Paul

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 6, 2020, 3:13pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/20 "2020-09-06T15:13:49Z")

</div>

(I moved this to Offtopic since the discussion is about construction regexs, not Julia code.)

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [September 6, 2020, 3:16pm UTC](https://discourse.julialang.org/t/regx-what-wrong/46059/21 "2020-09-06T15:16:50Z")

</div>

> [@programista](#):
>
> hex no! All data are just string UTF8

That… was not what I meant. I was asking if you considered `a` and `b` to be digits (as you were trying to match them with `\d`) because the numbers were in base 16 instead of base 10.
