# Regex capture next line in text file

**URL:** <https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361>\
**Category:** General Usage\
**Tags:** question, regex\
**Created:** [July 8, 2023, 7:28pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361 "2023-07-08T19:28:09Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Mulholland](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nick_mulholland/32/51373_2.png) [@Nick\_Mulholland](https://discourse.julialang.org/u/Nick_Mulholland)\
**Post date:** [July 8, 2023, 7:28pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/1 "2023-07-08T19:28:09Z")

</div>

I’m fairly next to Julia and was wondering if anyone had some advice. I’m loving the regex feature which is working great on a line by line query for example

```julia
open(filename) do file
        for (index, ln) in enumerate(eachline(file))
            m = match(r"\s+<minimizedAffinity>\s*",ln)
            if m != nothing
print(m)
end

```

However I struggling with how to do read in the whole file and find matches for a multiline pattern. i.e the float value that comes after the keyword.

```julia
(\s+<minimizedAffinity>[\n,\r]([-+]?[0-9]*\.?[0-9]+(e[-+]?[0-9]+)?))

```

the goal to extract a list of float values from the text file.

I can faff around with indexing the lines but there must be a more elegant solution?

---

<div class="post-metadata">

**Author:** ![gustaphe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gustaphe/32/18174_2.png) [@gustaphe](https://discourse.julialang.org/u/gustaphe)\
**Post date:** [July 8, 2023, 8:49pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/3 "2023-07-08T20:49:02Z")

</div>

How large is this file? If you can feasibly load the entire thing as a string I think you can just run that query (though you’ll want to remove the comma). Otherwise, you can do

```julia
function findnumber(file)
    lookfornumber = false
    for line in eachline(file)
        if lookfornumber
            m = match(<+just the number pattern+>, line)
            isnothing(m) || return m
            lookfornumber = false
        end
        m = match(<+just the preceding line pattern+>, line)
        if !isnothing(m)
            lookfornumber = true
        end
    end
    return nothing
end
m = open(findnumber, filename)

```

---

<div class="post-metadata">

**Author:** ![Nick\_Mulholland](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nick_mulholland/32/51373_2.png) [@Nick\_Mulholland](https://discourse.julialang.org/u/Nick_Mulholland)\
**Post date:** [July 8, 2023, 9:53pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/4 "2023-07-08T21:53:41Z")

</div>

Ah yes, that’s a clever way of doing it. Thanks you this worked!!

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [July 8, 2023, 11:10pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/5 "2023-07-08T23:10:43Z")

</div>

Because this pattern depends on two lines, the Iterators.Stateful trick can manage this logic. Try:

```julia
function findnumber2(file)
    itr = Iterators.Stateful(eachline(file))
    for line in itr
        m = match(r"Number:", line)
        if !isnothing(m)
            n = match(r"\d+", something(peek(itr), ""))
            isnothing(n) || return n
        end
    end
    return nothing
end

```

For example, for the above RegExs, the following file matches:

```julia
just
plain
324
63463
Number:
234234
fsdf
sdfg
Number:
sdfsf

```

and in Julia with the above function:

```julia
julia> testfile = IOBuffer("just\nplain\n324\n63463\nNumber:\n234234\nfsdf\nsdfg\nNumber:\nsdfsf")
IOBuffer(data=UInt8[...], readable=true, writable=false, seekable=true, append=false, size=59, maxsize=Inf, ptr=1, mark=-1)

julia> findnumber2(testfile)
RegexMatch("234234")

```

gustaphe’s answer `findnumber` finds this match as well.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 9, 2023, 12:18pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/6 "2023-07-09T12:18:15Z")

</div>

@Dan, why do we need `something(peek(itr),"")` and `peek(itr)` is not enough? Thank you.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [July 9, 2023, 1:35pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/7 "2023-07-09T13:35:51Z")

</div>

I suppose is to manage the case when the last line of the file is “Number:”

```julia
function findnumber21(file)
    itr = Iterators.Stateful(eachline(file))
    for line in itr
        m = match(r"Number:", line)
        if !isnothing(m)
            n = match(r"\d+", peek(itr))
            isnothing(n) || return n
        end
    end
    return nothing
end

julia> str="just\nplain\n324\n63463\nfsdf\nsdfg\nsdfsf\nNumber:"
"just\nplain\n324\n63463\nfsdf\nsdfg\nsdfsf\nNumber:"

julia> tf=IOBuffer(f)
IOBuffer(data=UInt8[...], readable=true, writable=false, seekable=true, append=false, size=61, maxsize=Inf, ptr=1, mark=-1)

julia> findnumber21(tf)
ERROR: MethodError: no method matching match(::Regex, ::Nothing)

Closest candidates are:
  match(::Regex, ::Union{SubString{String}, String}, ::Integer)
   @ Base regex.jl:374
  match(::Regex, ::Union{SubString{String}, String}, ::Integer, ::UInt32)
   @ Base regex.jl:374
  match(::Regex, ::InlineString, ::Integer)
   @ InlineStrings C:\Users\sprmn\.julia\packages\InlineStrings\rlLZO\src\InlineStrings.jl:711
  ...

Stacktrace:
 [1] findnumber21(file::IOBuffer)
   @ Main c:\Users\sprmn\.julia\environments\v1.9.0\regex_file.jl:36
 [2] top-level scope
   @ c:\Users\sprmn\.julia\environments\v1.9.0\regex_file.jl:46

julia> findnumber2(tf)

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [July 9, 2023, 3:53pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/8 "2023-07-09T15:53:19Z")

</div>

This avoids the need for the somethink() function (although it was nice to know about it).

```julia
using IterTools
function findnumber3(tf)
    itr = partition(eachline(tf),2,1)
    m=nothing
    for tl in itr
        m = isnothing(match(r"Number:", first(tl))) ? continue : match(r"\d+", last(tl))
        !isnothing(m) && return m
    end
end

```

I still have doubts about the need (apart from the usefulness in testing) of the Iterators.statefull() function

PS  
this formulation also holds for the more general problem where the lines to be inspected are k-positions apart

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [July 9, 2023, 4:45pm UTC](https://discourse.julialang.org/t/regex-capture-next-line-in-text-file/101361/9 "2023-07-09T16:45:00Z")

</div>

> [@gustaphe](#):
>
> How large is this file? If you can feasibly load the entire thing as a string I think you can just run that query (though you’ll want to remove the comma).

Even easier: `mmap` the file, wrap a [StringView](https://github.com/JuliaStrings/StringViews.jl) around it, and then run the desired multi-line regex query:

```julia
using StringViews, Mmap
open(filename, "r") do io
    s = StringView(mmap(io))
    for m in eachmatch(regex, s)
       # do something
    end
end

```

This way, the operating system will page the file into memory as needed, even if it is huge, but it will still _act_ like a string you loaded all at once.

As a bonus, this will also avoid the `eachline` performance cost of allocating a new string for every line. (Though this can be done more efficiently with the [ViewReader.jl package](https://github.com/rickbeeloo/ViewReader) and with the upcoming [`copyuntil` in Julia 1.11](https://github.com/JuliaLang/julia/pull/48273).)
