# Reading data text files delimited with both spaces & tabs

**URL:** <https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851>\
**Category:** General Usage\
**Tags:** csv, io\
**Created:** [July 18, 2021, 10:49am UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851 "2021-07-18T10:49:23Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 10:49am UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/1 "2021-07-18T10:49:23Z")

</div>

Is it possible in CSV.jl, DelimitedFiles.jl, or other, to read text files that contain **both** spaces & tabs, in the header and/or data sections?

An example is provided below.

NB: _the hidden spaces and tabs should be there after copy and paste, also added row with missing value_

```julia
Col1 Col2 Col3 
1	1 012
2 1 013
2 1   
2 1 015

```

Currently, the workaround used is to replace all tabs by spaces in a text editor and then read the file using CSV.jl:

```julia
using CSV, DataFrames
df = CSV.read(file, header=1, delim=" ", ignorerepeated=true, type=Int64, DataFrame)

```

Thank you.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 18, 2021, 11:17am UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/2 "2021-07-18T11:17:26Z")

</div>

`readdlm` does so by default?

```julia
julia> using DelimitedFiles
       readdlm("/tmp/testfile.txt")
3×3 Matrix{Any}:
  "Col1" "Col2" "Col3"
 1 1 12
 2 1 13

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 11:54am UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/3 "2021-07-18T11:54:19Z")

</div>

@GunnarFarneback, thanks for your response. The result is a matrix of Any that needs further parsing in order to use the convenience tools in DataFrames and DataFramesMeta to handle missings, etc. (_my input example was too simple_).

The following code (adapted from [here](https://discourse.julialang.org/t/converting-a-matrix-to-a-dataframe/6114/4)) converts the matrix of Any into a dataframe, allowing further processing as indicated.

It does not seem to be as “efficient” as CSV.jl (_where there is a lot of machinery for handling types, etc._) but it works and will mark your response as a solution, if no other ideas pop up.

```julia
using DelimitedFiles
M = readdlm(file)
header = M[1,:]
data = M[2:end,:]
df = DataFrame(data,:auto)
rename!(df, names(df) .=> Symbol.(header[:]))

```

Thanks.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 18, 2021, 12:34pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/4 "2021-07-18T12:34:05Z")

</div>

You can do the replacement in Julia without too much fuzz:

```julia
df = CSV.read(IOBuffer(replace(read("/tmp/testfile.txt"), UInt8('\t') => UInt8(' '))), header=1, delim=" ", ignorerepeated=true, type=Int64, DataFrame)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 12:47pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/5 "2021-07-18T12:47:11Z")

</div>

Gunnar, this is really brilliant.

No clue how you managed to implement the replacement operation using the IOBuffer.

The docs mention such intermediate operations only vaguely. Your example should be part of the docs, IMHO.

Thanks again.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 18, 2021, 1:02pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/6 "2021-07-18T13:02:30Z")

</div>

There’s nothing deep going on here. If split into intermediate results:

```julia
a = read(file); # Read a full file into a UInt8 vector
b = replace(a, UInt8('\t') => UInt8(' ')) # Replace tabs by spaces
c = IOBuffer(b) # IOBuffer may optionally operate on a pre-existing array

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 1:08pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/7 "2021-07-18T13:08:32Z")

</div>

So, the file is actually read “twice”? Once into IOBuffer memory and then again from the memory buffer with CSV.read()?

---

<div class="post-metadata">

**Author:** ![cjdoris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cjdoris/32/213133_2.png) [@cjdoris](https://discourse.julialang.org/u/cjdoris)\
**Post date:** [July 18, 2021, 3:16pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/8 "2021-07-18T15:16:30Z")

</div>

Unless you’re sure the file is ASCII, you should probably do

```julia
a = read(file, String)
b = replace(a, '\t' => ' ')
c = IOBuffer(b)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 3:20pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/9 "2021-07-18T15:20:32Z")

</div>

@cjdoris, how do you plug-it in the one-liner above?  
Get error:

```julia
ERROR: MethodError: no method matching findnext(::UInt8, ::String, ::Int64)

```

---

<div class="post-metadata">

**Author:** ![cjdoris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cjdoris/32/213133_2.png) [@cjdoris](https://discourse.julialang.org/u/cjdoris)\
**Post date:** [July 18, 2021, 3:25pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/10 "2021-07-18T15:25:57Z")

</div>

Can you show us what you tried?

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [July 18, 2021, 3:36pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/11 "2021-07-18T15:36:14Z")

</div>

It reads it more than two times, since `replace` also generate intermediate array.  
If you want to avoid it, then you should better use `map!`

```julia
a = read(file)
map!(c -> c == UInt8('\t') ? UInt8(' ') : c, a, a)
df = CSV.read(IOBuffer(a), header=1, delim=" ", ignorerepeated=true, type=Int64, DataFrame)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 4:11pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/12 "2021-07-18T16:11:34Z")

</div>

@cjdoris, sorry messed it up here. It works perfectly 🙂

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 18, 2021, 4:28pm UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/13 "2021-07-18T16:28:35Z")

</div>

> [@Skoffer](#):
>
> It reads it more than two times, since `replace` also generate intermediate array

Thank you, very good to know.

Should we use piping to adapt your solution to Gunnar’s one-liner?  
E.g.:

```julia
df = CSV.read(IOBuffer(read(file) |> x->map!(c-> c==UInt8('\t') ? UInt8(' ') : c, x, x)), header=1, delim=" ", ignorerepeated=true, type=Int64, DataFrame)

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 19, 2021, 4:54am UTC](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851/14 "2021-07-19T04:54:14Z")

</div>

I also saw such problems that need a simple preprocessing of a file before feeding it to CSV or another reader library, and also had to resort to similar workarounds.  
It should be possible to implement general functions like `map_stream_bytes(::Function, io) -> io`, `map_stream_chars`, `map_stream_lines` (and `filter_...` as well): they would take and return IO streams without reading all the content to memory at once. Or maybe something like `iostream_to_iterator_of_chars`/`iostream_from_iterator_of_chars` is even better.  
Didn’t go as far as implementing these functions myself though.
