# Read text file containing proper CSV data chunks

**URL:** <https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828>\
**Category:** Data\
**Tags:** dataframes, csv\
**Created:** [December 19, 2023, 8:24pm UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828 "2023-12-19T20:24:16Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 19, 2023, 8:24pm UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/1 "2023-12-19T20:24:16Z")

</div>

I have a text file that contains comma-separated entries, and I am interested only in the rows that begin with the character `A`:

```julia
file = """
C comment-1
B,9,5.5,20:30:33
A,2,1.5,20:31:15
C comment-2
A,0,0.5,22:57:00
C comment-3
"""
str = filter(x -> first(x) == 'A', readlines(IOBuffer(file)))

```

The above command extracts the rows of interest.

How can we streamline this using CSV.jl to read the filtered rows of interest?

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [December 20, 2023, 12:01am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/2 "2023-12-20T00:01:28Z")

</div>

Maybe this?

```julia-repl
julia> CSV.read(IOBuffer(join(str, "\n")), DataFrame, header=false)
2×4 DataFrame
 Row │ Column1 Column2 Column3 Column4  
     │ String1 Int64 Float64 Time     
─────┼─────────────────────────────────────
   1 │ A 2 1.5 20:31:15
   2 │ A 0 0.5 22:57:00

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 20, 2023, 12:10am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/3 "2023-12-20T00:10:27Z")

</div>

Thanks @lmiq, that was nice!

I would also be interested in a “row-by-row” CSV.read solution, if such a possibility exists.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 20, 2023, 12:44am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/4 "2023-12-20T00:44:50Z")

</div>

I implemented something like that in FamaFrenchData.jl

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 20, 2023, 7:54am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/5 "2023-12-20T07:54:45Z")

</div>

Thanks @tbeason. After a quick look at your code, it appears to read all the lines in memory, identify the blocks of interest, then perform CSV.read by block followed by a merge.

What I was asking for was a row-by-row CSV.jl reading with a filter to process only the lines of interest.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 20, 2023, 8:51am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/6 "2023-12-20T08:51:53Z")

</div>

You can use [`CSV.rows`](https://csv.juliadata.org/stable/reading.html#CSV.Rows) but it’s less efficient and convenient than `CSV.read`.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 20, 2023, 8:59am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/7 "2023-12-20T08:59:00Z")

</div>

There aren’t many examples of using CSV.Rows, do you know of any that could be a source of inspiration?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 20, 2023, 10:58am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/8 "2023-12-20T10:58:26Z")

</div>

No. :-/

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [December 20, 2023, 11:01am UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/9 "2023-12-20T11:01:24Z")

</div>

> [@rafael.guerra](#):
>
> ```julia
> file = """
> C comment-1
> B,9,5.5,20:30:33
> A,2,1.5,20:31:15
> C comment-2
> A,0,0.5,22:57:00
> C comment-3
> """
> 
> ```

You can do it by hand, by pushing to the DataFrame:

```julia-repl
julia> function read_df(file)
           df = nothing
           for line in readlines(IOBuffer(file))
       	if line[1] == 'A'
                   if isnothing(df)
       		    df = CSV.read(IOBuffer(line), DataFrame, header=false)
                   else
                       push!(df, CSV.read(IOBuffer(line), DataFrame, header=false)[1,:])
                   end
               end
           end
           return df
       end
read_df (generic function with 2 methods)

julia> read_df(file)
2×4 DataFrame
 Row │ Column1 Column2 Column3 Column4  
     │ String1 Int64 Float64 Time     
─────┼─────────────────────────────────────
   1 │ A 2 1.5 20:31:15
   2 │ A 0 0.5 22:57:00

```

(sorry for the badly indented code, my terminal went crazy)

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 20, 2023, 12:34pm UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/10 "2023-12-20T12:34:35Z")

</div>

Thanks @lmiq, it looks good but not as simple as one might have hoped.

The following is a bit shorter but still not ideal:

```julia
cnames = [:Code, :Counter, :Value, :Time]
types = [String[], Int[], Float64[], Time[]] 
df = DataFrame([name => type for (name,type) in zip(cnames,types)])
for line in readlines(IOBuffer(file))
    line[1] == 'A' && append!(df, CSV.read(IOBuffer(line), DataFrame, header=cnames))
end

```

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 20, 2023, 2:27pm UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/11 "2023-12-20T14:27:26Z")

</div>

Ah I see, I misunderstood. Sorry for the misdirection.

CSV.jl does have a `comment` argument. Perhaps it could be faster to read all the data, except the comment lines, and then use `subset` to filter out all the lines that do not start with `A`? Obviously the memory footprint will be a little larger than a line-by-line method but by reading the entire file you can leverage the multithreaded and optimized capabilities of CSV.jl.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 20, 2023, 3:00pm UTC](https://discourse.julialang.org/t/read-text-file-containing-proper-csv-data-chunks/107828/12 "2023-12-20T15:00:55Z")

</div>

> [@tbeason](#):
>
> [CSV.jl](https://juliahub.com/ui/Packages/CSV) does have a `comment` argument.

This works great but only filters ‘C’, what one would need is `comment!='A'`
