# Split .txt file into several DataFrames

**URL:** <https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526>\
**Category:** General Usage\
**Tags:** dataframes, csv\
**Created:** [March 23, 2023, 6:23pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526 "2023-03-23T18:23:18Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![fdekerme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fdekerme/32/43574_2.png) [@fdekerme](https://discourse.julialang.org/u/fdekerme)\
**Post date:** [March 23, 2023, 6:23pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/1 "2023-03-23T18:23:18Z")

</div>

Hello 😁,  
I would like to load a .txt file and split it inot several DataFrames .  
The file has the following general structure:

```julia
#Name:
#Id:
#Dosename:
#RoiName:brain
#Roi volume
#Unit: Gy
0.000 100.000
0.290 89.0
0.580 67.8
0.870 55.0
1.161 43.1
1.451 21.3

#RoiName:neck
#River volume 
#Unit: Gy
0.000 100.000
0.081 89.1
0.162 68.3
0.243 56.9

```

The idea would be to split the file at each “`#RoiName:`” and make a several DataFrames :

> brain\_df =  
> x y  
> 0.000 100.000  
> 0.290 89.0  
> 0.580 67.8  
> 0.870 55.0  
> 1.161 43.1  
> 1.451 21.3
> 
> neck\_df =  
> x y  
> 0.000 100.000  
> 0.081 89.1  
> 0.162 68.3  
> 0.243 56.9

I tried to load my .txt file with CSV.jl and then to convert it with DataFrames.jl, but no idea to “split” it has described above.

Thanks in advance ! 😁

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 23, 2023, 6:51pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/2 "2023-03-23T18:51:19Z")

</div>

```julia
using DataFrames
using CSV
input = """
#Name:
#Id:
#Dosename:
#RoiName:brain
#Roi volume
#Unit: Gy
0.000 100.000
0.290 89.0
0.580 67.8
0.870 55.0
1.161 43.1
1.451 21.3

#RoiName:neck
#River volume 
#Unit: Gy
0.000 100.000
0.081 89.1
0.162 68.3
0.243 56.9
"""
io = IOBuffer(input)
dfs = DataFrame[]
buffer = String[]
for line in eachline(io)
    if !startswith(line, "#")
        if isempty(line)
            if !isempty(buffer)
                push!(dfs, CSV.read(IOBuffer(join(buffer, "\n")), DataFrame, header=["x", "y"]))
                empty!(buffer)
            end
        else
            push!(buffer, line)
        end
    end
end
if !isempty(buffer)
    push!(dfs, CSV.read(IOBuffer(join(buffer, "\n")), DataFrame, header=["x", "y"]))
    empty!(buffer)
end

```

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [March 23, 2023, 7:00pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/3 "2023-03-23T19:00:13Z")

</div>

Alternatively, using `readuntil`:

```julia
julia> open("sample.txt") do f
           df = Dict{String,DataFrame}()
           readuntil(f, "#RoiName:")
           while !eof(f)
               df[readline(f)] = CSV.read(IOBuffer(readuntil(f, "#RoiName:")), DataFrame; 
                   comment="#", header=["x", "y"])
           end
           df
       end

Dict{String, DataFrame} with 2 entries:
  "brain" => 6×2 DataFrame…
  "neck" => 4×2 DataFrame…

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [March 23, 2023, 7:38pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/4 "2023-03-23T19:38:28Z")

</div>

@stillyslalom, to run your nice code in Julia 1.8.5, I need to split the inner loop assignment as follows:

```julia
str = readline(f)
df[str] = ...

```

Is this a new feature in Julia 1.9?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [March 23, 2023, 7:40pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/5 "2023-03-23T19:40:52Z")

</div>

Nope, just an erroneous simplification on my part - the LHS gets evaluated after the RHS. This works:

```julia
julia> open("sample.txt") do f
                  df = Dict{String,DataFrame}()
                  readuntil(f, "#RoiName:")
                  while !eof(f)
                      name, rest = readline(f), IOBuffer(readuntil(f, "#RoiName:"))
                      df[name] = CSV.read(rest, DataFrame; comment="#", header=["x", "y"])
                  end
                  df
              end
Dict{String, DataFrame} with 2 entries:
  "brain" => 6×2 DataFrame…
  "neck" => 4×2 DataFrame…

```

---

<div class="post-metadata">

**Author:** ![fdekerme](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fdekerme/32/43574_2.png) [@fdekerme](https://discourse.julialang.org/u/fdekerme)\
**Post date:** [March 27, 2023, 4:10pm UTC](https://discourse.julialang.org/t/split-txt-file-into-several-dataframes/96526/6 "2023-03-27T16:10:02Z")

</div>

Thank you very much for this solution!
