# Reading a file from line x to line y

**URL:** <https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363>\
**Category:** General Usage\
**Tags:** csv\
**Created:** [May 16, 2024, 3:08pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363 "2024-05-16T15:08:37Z")\
**Posts on this page:** 9\
**Page:** 2

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 18, 2024, 10:08am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/22 "2024-05-18T10:08:01Z")

</div>

Thanks all ! Trying various suggestions.

About the mmap version 1, I tried this:

```julia
using Mmap

const input_file="file.txt"
const output_file="out.txt"
    
read_start=2
read_stop=4

function process_file(in_fn, startline, stopline)
    f = open(in_fn, "r")
    mm = Mmap.mmap(f, Vector{UInt8})
    l = 1
    pos = 0
    while l < startline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    startpos = pos
    while l <= stopline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    stoppos = pos
    #write(stdout, @view mm[startpos:stoppos])
    open(output_file, "w") do file
        write(file, @view mm[startpos:stoppos])
    end
    close(f)
end

process_file(input_file, read_start, read_stop)

```

But while the result of the cut is ok, there’s a blank line at the top of the out.txt file. I don’t get why.

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 18, 2024, 10:42am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/23 "2024-05-18T10:42:39Z")

</div>

About the Julian version, I did this:

```julia
using Mmap

const input_file="file.txt"
const output_file="out.txt"

read_start=2
read_stop=4

function process_file(in_fn, startline, stopline)
    open(in_fn, "r") do f
        mm = Mmap.mmap(f, Vector{UInt8})
        skipline = let mm = mm
            (pos, l) -> last(findnext([UInt8('\n')], mm, pos+1))
        end
        startpos = foldl(skipline, 1:startline; init=1)+1
        stoppos = foldl(skipline, startline:stopline; init=startpos)
        lines=readlines(IOBuffer(@view mm[startpos:stoppos]))
        open(output_file, "w") do io
            for ll = lines
                write(io, ll * "\n")
            end
        end
    end
end

process_file(input_file, read_start, read_stop)

```

It’s a bit slower, and the range is not ok as I got:

```julia
3 ccc
4 ddd
5 eee

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 18, 2024, 1:07pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/24 "2024-05-18T13:07:24Z")

</div>

To fix first version:

```julia
function process_file(in_fn, startline, stopline)
... // fix following line
        write(file, @view mm[startpos+1:stoppos-1])
...
end

```

[/quote]

To fix second version:

```julia
function process_file(in_fn, startline, stopline)
... // fix following lines:
        startpos = foldl(skipline, 1:startline-1; init=1)+1 
        stoppos = foldl(skipline, startline:stopline; init=startpos)-1
...
end

```

These are just little one-off errors. Easy to fix. The overall performance beats `sed` (and other parsing solutions).

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 18, 2024, 3:26pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/25 "2024-05-18T15:26:49Z")

</div>

Thanks al lot !  
Here’s the corrected versions. Note that I feel safer with a \newline at the end of the last line.

First version:

```julia
using Mmap

const input_file="file.txt"
const output_file="out.txt"
    
read_start=2
read_stop=4

function process_file(in_fn, startline, stopline)
    f = open(in_fn, "r")
    mm = Mmap.mmap(f, Vector{UInt8})
    l = 1
    pos = 0
    while l < startline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    startpos = pos
    while l <= stopline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    stoppos = pos
    #write(stdout, @view mm[startpos:stoppos])
    open(output_file, "w") do file
        write(file, @view mm[startpos+1:stoppos])
    end
    close(f)
end

process_file(input_file, read_start, read_stop)

```

Second version:

```julia
using Mmap

const input_file="file.txt"
const output_file="out.txt"

read_start=2
read_stop=4

function process_file(in_fn, startline, stopline)
    open(in_fn, "r") do f
        mm = Mmap.mmap(f, Vector{UInt8})
        skipline = let mm = mm
            (pos, l) -> last(findnext([UInt8('\n')], mm, pos+1))
        end
        startpos = foldl(skipline, 1:startline-1; init=1)+1
        stoppos = foldl(skipline, startline:stopline; init=startpos)
        lines=readlines(IOBuffer(@view mm[startpos:stoppos]))
        open(output_file, "w") do io
            for ll = lines
                write(io, ll * "\n")
            end
        end
    end
end

process_file(input_file, read_start, read_stop)

```

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 18, 2024, 3:48pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/26 "2024-05-18T15:48:50Z")

</div>

Meanwhile I tested a way to load the extract into an array structure for further processings:

```julia
using Mmap
using DelimitedFiles

const input_file="file.txt"
    
read_start=2
read_stop=4

function process_file(in_fn, startline, stopline)
    f = open(in_fn, "r")
    mm = Mmap.mmap(f, Vector{UInt8})
    l = 1
    pos = 0
    while l < startline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    startpos = pos
    while l <= stopline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    stoppos = pos
    csv=readdlm(mm[startpos:stoppos]; use_mmap=true)
    close(f)
    return csv
end

csv_like=process_file(input_file, read_start, read_stop)
println(csv_like)
println(csv_like[1,2])

```

Output:

```julia
Any[2 "bbb"; 3 "ccc"; 4 "ddd"]
bbb

```

It’s working well.

---

<div class="post-metadata">

**Author:** ![Eben60](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eben60/32/13475_2.png) [@Eben60](https://discourse.julialang.org/u/Eben60)\
**Post date:** [May 18, 2024, 10:00pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/27 "2024-05-18T22:00:29Z")

</div>

After you have found the starting position in the file, you could also do:

```julia
types = [DateTime, Float64, Float64, Float64]
fmt = dateformat"yyyymmddHHMMSS"
f = open(input_file, "r")
seek(f, startpos)
df = DataFrame(CSV.File(f; limit=yourlimit, types, dateformat=fmt)) # untested!

```

`CSV` provides options to parse the input on read, and `DataFrames` is _the_ tool for the analysis of that kind of data (surely, all is a matter of personal preference).

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 22, 2024, 11:30am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/28 "2024-05-22T11:30:19Z")

</div>

I encounter a problem when the extract starts at the first line of the file:

```julia
BoundsError: attempt to access 22218-element Vector{UInt8} at index [0:16842]

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 22, 2024, 11:48am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/29 "2024-05-22T11:48:33Z")

</div>

There must be an off-by-one error. Try initializing `pos = 1` instead of `pos = 0`.

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 22, 2024, 12:28pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/30 "2024-05-22T12:28:00Z")

</div>

Thanks, it works !

I wanted to edit my previous codes, but I don’t find how to edit them…

[Previous page](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363.md?page=1)
