# Reading a file from line x to line y

**URL:** <https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363>\
**Category:** General Usage\
**Tags:** csv\
**Created:** [May 16, 2024, 3:08pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363 "2024-05-16T15:08:37Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 16, 2024, 3:08pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/1 "2024-05-16T15:08:37Z")

</div>

Hi,

I would like to read a portion of a lot of big files, as I know that the lines that interest me are located within line x and line y  
Currently, I load the files with CSV.File and read parts, but the loading overhead is way too much.  
The same process in shell (with sed/grep…) takes 27 minutes, Julia takes 139.  
Should I call grep from Julia ?

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [May 16, 2024, 3:15pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/2 "2024-05-16T15:15:34Z")

</div>

Did you try these?

[skipto](https://csv.juliadata.org/stable/reading.html#skipto)  
[limit](https://csv.juliadata.org/stable/reading.html#limit)

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 17, 2024, 9:23am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/3 "2024-05-17T09:23:03Z")

</div>

Unfortunately, as far as I’ve tested it, skipto/limit has poor performance:

```julia
@time csv=CSV.File("data/big_file.txt"; header=false, skipto=1501245, limit=450)
  5.636217 seconds (137.98 M allocations: 2.146 GiB, 0.36% gc time, 0.08% compilation time)
450-element CSV.File:
...

```

In shell:

```julia
time sed -n '1501245,1501694p;1501695q' data/big_file.txt | wc -l
450

real	0m0,061s

```

---

<div class="post-metadata">

**Author:** ![HanD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hand/32/213908_2.png) [@HanD](https://discourse.julialang.org/u/HanD)\
**Post date:** [May 17, 2024, 9:57am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/4 "2024-05-17T09:57:57Z")

</div>

You could try using `eachline`, `Iterators.drop` and `Iterators.take` to assemble the lines you need into an in-memory stream, and then read the data from there. Depending on the amount of data you need to scan, you can either join the lines in question into a single string in memory, or use an `IOBuffer` and an asynchronous task to feed the buffer.

But you made me wonder… If just grepping the CSV files takes 27 minutes, wouldn’t you be better off by importing the CSV data into an SQL database once, and then fetching the necessary data using queries instead? You could use SQLite or DuckDB, or any SQL solution of your choice…

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 17, 2024, 10:04am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/5 "2024-05-17T10:04:44Z")

</div>

Looks like very complicated for such a simple operation.

For various reasons, I can’t use SQL…

I’m currently experimenting with:

```julia
block=readchomp(`sed -n "$start,$end p;$end_plus_one q" $filename`)

```

It’s fast but it doesn’t return an array.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [May 17, 2024, 10:58am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/6 "2024-05-17T10:58:03Z")

</div>

Please check also [this thread](https://discourse.julialang.org/t/skipping-a-lot-of-lines-in-csv-read-allocates-too-much-memory/110121).

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [May 17, 2024, 11:07am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/7 "2024-05-17T11:07:12Z")

</div>

Only ever used Windows so take all of this with a grain of salt. Looking it up, sed is a stream editor working on one line at a time, wc -l counts lines, and piped processes execute concurrently (maybe like broadcasting fusing into 1 loop), so my guess is that shell command increments a count for each line read. By contrast, `CSV.File` must process lines into a tabular structure with a fixed number of named columns, which is a lot of work to find out. `countlines(...)` runs much faster than `length(CSV.File(...))`, at least on the small files I have lying around. My hunch is you don’t want all that processing, you just need to read lines of text.

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 17, 2024, 11:16am UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/8 "2024-05-17T11:16:48Z")

</div>

If the file contains a big dataset, usually the content of the lines has data identifying the data. In this case, random access and binary search to find the initial position to read and stop condition could blow the other methods out of the water.

For example, if each line contains a sequential ID, or any sorted ID (for which you know the ID of the initial and final lines).

The TL;DR: If data is sorted use binary-search.

---

<div class="post-metadata">

**Author:** ![HanD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hand/32/213908_2.png) [@HanD](https://discourse.julialang.org/u/HanD)\
**Post date:** [May 17, 2024, 1:15pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/9 "2024-05-17T13:15:11Z")

</div>

> [@Ju\_ska](#):
>
> For various reasons, I can’t use SQL…

Well, that’s a bit vague. I understand if you can’t _create_ intermediate SQL files, but it’s odd if you can’t use SQL _at all_…

DuckDB, for instance, supports querying directly from a CSV file. I don’t know the performance of this operation, but generally, DuckDB is quite performant. So maybe you can have a look at this: [CSV Import – DuckDB](https://duckdb.org/docs/data/csv/overview.html), and then you can narrow the query using SQL syntax. The Julia library for DuckDB also supports importing the data into e.g., `DataFrames`, so the result should be compatible with what you get from the CSV lib.

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 17, 2024, 2:03pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/10 "2024-05-17T14:03:22Z")

</div>

I’d like to keep things simple for something as simple as this…  
For now I crafted a run command with sed inside, and write the extract in a ram folder, then I read the extract as a csv. I now run under 18 minutes \o/ .

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 17, 2024, 4:19pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/11 "2024-05-17T16:19:07Z")

</div>

Can you provide a few lines from the `big_file` to have as a sample, and what is the size in MB and in lines of the file?  
(for each problem there are specific optimizations)

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 17, 2024, 4:30pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/12 "2024-05-17T16:30:10Z")

</div>

It’s data from Earth magnetic field, reshaped under this form:

```julia
20200320201739 29204.51 5900.46 -18887.12
20200320201740 29204.53 5900.45 -18887.09
20200320201741 29204.54 5900.46 -18887.11
20200320201742 29204.54 5900.45 -18887.10
20200320201743 29204.58 5900.45 -18887.11
20200320201744 29204.59 5900.47 -18887.08
20200320201745 29204.60 5900.49 -18887.10
20200320201746 29204.63 5900.49 -18887.09

```

Files can be around GiB size.

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 17, 2024, 5:05pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/14 "2024-05-17T17:05:17Z")

</div>

Well, this exactly provides the kind of specific additional info which could solve the speed problem:

1. The file is ASCII one byte per codepoint.
2. Lines seem to be of about equal length and tame ASCII letters.
3. Data in each line provides _sorted_ indication of which line it is, and the required lines can probably be devined from data content (allowing random-access binary-search).

The points (1)-(3) are not working code, but getting some 500 lines from the middle of such a file, should be doable and probably even much faster than `sed`.

Regarding point (3), another bit of interesting info is how the chosen range of lines is determined (is there an “index” file which specifies line numbers)?

For example, a vectorized loop counting the number of `UInt('\n')` can very quickly scan an `mmap`ed file to find the right line, even without random-access directly to mid-file.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 17, 2024, 5:30pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/15 "2024-05-17T17:30:55Z")

</div>

Looks like I had forgotten to add one necessary option to do this. With GMT.jl (master version for now)

```julia
using GMT

D = gmtread("ppt20240301vsec.sec", data=true, inrows="100:200")
Attribute table
┌─────────┐
│ Timecol │
├─────────┤
│ 1 │
└─────────┘
BoundingBox: [1.7092512e9, 1.7092512e9, 0.020555555555555556, 0.04833333333333333, 61.0, 61.0, 29136.36, 29137.42, 5945.51, 5945.67, -18656.91, -18656.72, 35102.42, 88888.0]

101×7 GMTdataset{Float64, 2}
 Row │ Time col.2 col.3 col.4 col.5 col.6 col.7
─────┼────────────────────────────────────────────────────────────────────────
   1 │ 2024-03-01T00:01:14 0.0 61.0 29137.4 5945.56 -18656.8 88888.0
   2 │ 2024-03-01T00:01:15 0.0 61.0 29137.4 5945.53 -18656.8 35103.3
   3 │ 2024-03-01T00:01:16 0.0 61.0 29137.4 5945.58 -18656.8 88888.0
   4 │ 2024-03-01T00:01:17 0.0 61.0 29137.4 5945.6 -18656.8 88888.0
   5 │ 2024-03-01T00:01:18 0.0 61.0 29137.4 5945.63 -18656.8 88888.0

```

Note, GMT checks lots of things when reading text tables so this may not be the fastest of the solutions.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 17, 2024, 5:35pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/16 "2024-05-17T17:35:27Z")

</div>

Dan, from another thread, my guess is that first column is time. Number of seconds since ???

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 17, 2024, 5:40pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/17 "2024-05-17T17:40:53Z")

</div>

```julia
20200320201743 29204.58 5900.45 -18887.11

```

`20200320201743` is a DateTime `2020`-`03`-`20`-`20:17:43`

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [May 17, 2024, 5:50pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/18 "2024-05-17T17:50:10Z")

</div>

Did you try to tell CSV the types of the columns?

`types`: a single `Type`, `AbstractVector` or `AbstractDict` of types, or a function of the form `(i, name) -> Union{T, Nothing}` to be used for column types; if a single `Type` is provided, _all_ columns will be parsed with that single type; an `AbstractDict` can map column index `Integer`, or name `Symbol` or `String` to type for a column, i.e. `Dict(1=>Float64)` will set the first column as a `Float64`, `Dict(:column1=>Float64)` will set the column named `column1` to `Float64` and, `Dict("column1"=>Float64)` will set the `column1` to `Float64`; if a `Vector` is provided, it must match the # of columns provided or detected in `header`. If a function is provided, it takes a column index and name as arguments, and should return the desired column type for the column, or `nothing` to signal the column’s type should be detected while parsing.

Did you try to read the file using multiple threads?

> **[Reading · CSV.jl](https://csv.juliadata.org/stable/reading.html#CSV.File)**
>
> Documentation for CSV.jl.

Did you try InMemoryDatasets.jl ? This function reads a csv file: [Reading and Writing · DLMReader](https://docs.juliahub.com/General/DLMReader/stable/man/read/)

In my tests it was consistently faster than CSV.jl …

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [May 17, 2024, 10:19pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/19 "2024-05-17T22:19:10Z")

</div>

To concretize some on my earlier suggestions, here is code to print lines from x to y:

```julia
using Mmap

function process_file(in_fn, startline, stopline)
    f = open(in_fn, "r")
    mm = Mmap.mmap(f, Vector{UInt8})
    l = 1
    pos = 0
    while l < startline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    startpos = pos
    while l <= stopline
        pos = last(findnext([UInt8('\n')], mm, pos+1))
        l += 1
    end
    stoppos = pos
    write(stdout, @view mm[startpos:stoppos])
    close(f)
end

```

Using this code, to replicate the lines from the first post:

```julia
julia> @time process_file("big_file.txt", 1501245,1501694)

20200320201743 29204.58 5900.45 -18887.11
:
:
20200320201744 29204.59 5900.47 -18887.08
  0.036912 seconds (1.50 M allocations: 91.657 MiB, 10.68% gc time)

```

which takes 37ms vs 59ms of the `sed` command:

```julia
$ time sed -n '1501245,1501694p;1501695q' big_file.txt
:
real	0m0.059s

```

As usual, different methods converge to same orders of magnitude.

UPDATE: A more Julian version of the above function (taking the same amount of time) returns a vector of line strings:

```julia
function process_file(in_fn, startline, stopline)
    open(in_fn, "r") do f
        mm = Mmap.mmap(f, Vector{UInt8})
        skipline = let mm = mm
            (pos, l) -> last(findnext([UInt8('\n')], mm, pos+1))
        end
        startpos = foldl(skipline, 1:startline; init=1)+1
        stoppos = foldl(skipline, startline:stopline; init=startpos)
        return readlines(IOBuffer(@view mm[startpos:stoppos]))
    end
end

```

---

<div class="post-metadata">

**Author:** ![technocrat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/technocrat/32/220947_2.png) [@technocrat](https://discourse.julialang.org/u/technocrat)\
**Post date:** [May 17, 2024, 10:21pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/20 "2024-05-17T22:21:17Z")

</div>

Have you tried `CSV.Chunks`?

```
• skipto::Integer: specifies the row where the data starts in the csv file;
   by default, the next row after the header row(s) is used. If header=0, then
   the 1st row is assumed to be the start of data; providing a skipto argument
   does not affect the header argument. Note that if a row number skipto and
   comment or ignoreemptyrows are provided, the data row will be the first
   non-commented/non-empty row after the row number, meaning if the provided
   row number is a commented row, the data row will actually be the next
   non-commented row.

• footerskip::Integer: number of rows at the end of a file to skip parsing.
   Do note that commented rows (see the comment keyword argument) do not count
   towards the row number provided for footerskip, they are completely ignored
   by the parser

```

---

<div class="post-metadata">

**Author:** ![Ju\_ska](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Ju\_ska](https://discourse.julialang.org/u/Ju_ska)\
**Post date:** [May 17, 2024, 10:49pm UTC](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363/21 "2024-05-17T22:49:28Z")

</div>

Thanks, I’ll check that.  
For now I use:

file.txt

```julia
1 aaa
2 bbb
3 ccc
4 ddd
5 eee
6 fff

```

Code

```julia
input_file="file.txt"
output_file="out.txt"

read_start=2
read_stop=4
quit=read_stop + 1

run(`/bin/sh -c " sed -n \"$read_start, $read_stop p; $quit q\" $input_file > $output_file"`)

```

Result

```julia
2 bbb
3 ccc
4 ddd

```

[Next page](https://discourse.julialang.org/t/reading-a-file-from-line-x-to-line-y/114363.md?page=2)
