# Reading text data: \`readdlm\` is deprecated, so how CSV package is used?

**URL:** <https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955>\
**Category:** General Usage\
**Tags:** question, csv, io, delimitedfiles, text-data\
**Created:** [July 12, 2024, 9:15am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955 "2024-07-12T09:15:11Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![ryofurue](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ryofurue/32/24531_2.png) [@ryofurue](https://discourse.julialang.org/u/ryofurue)\
**Post date:** [July 12, 2024, 9:15am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/1 "2024-07-12T09:15:11Z")

</div>

In the issue tracker of `DelimitedFiles`

> <https://github.com/JuliaData/DelimitedFiles.jl/issues/1>
>
> readdlm ignores all-white spaces by default when a delimiter is not specified. H…owever, when one wants to specify the data type to be read it is obligatory to specify the delimiter too...
> 
> \`readdlm(source, delim::AbstractChar, T::Type, eol::AbstractChar; header=false, skipstart=0, skipblanks=true, use\_mmap, quotes=true, dims, comments=false, comment\_char='#')\`
> 
> Then, in the following case, 
> 
> \`readdlm(file, ' ', Float64, comments=true)\`
> 
> the function doesn't ignore the initial whitespace because the delimiter is \`' '\`, only 1 whitespace. Then the program crashes with for example
> 
> \`\`\`
> 2 3
> 1 3
> \`\`\`
> There should be a flag to ignore all chars that match with the delimiter or just be able to specify the type like this
> 
> \`readdlm(file, type=Float64, comments=true)\`
> 
> however this brings the problem that if the delimiter is not a whitespace the problem will persist.

I found a comment

> `readdlm` is also effectively deprecated and the CSV package should be used

about reading text data files.

I’m trying to read a plain text file listing numbers delimited by newlines and spaces.

Does somebody know where to find a simple tutorial to use the CSV package to do this?

I just blindly tried the CSV package. I got an object containing objects like `Vector{CSV.Row}` and I didn’t know how to examine whether my read was successful or not.

The data file includes comment lines indicated by “#”. The file can be read with

```julia
using DelimitedFiles
a = readdlm("my-datafile.txt"; comments=true)
a[2,5] # -> the number at row 2, column 5

```

This is a very simple interface. You just get a 2D array. How does one use the CSV package to do this?

Aside: The data file I’m trying to read turns out including the unicode “zero width no-break space”, which `readdlm()` fails to handle. So, I was trying to do something about it, when I found the deprecation comment above.

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 12, 2024, 10:12am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/2 "2024-07-12T10:12:58Z")

</div>

The first step to try blindly is probably this:

```julia
using CSV, DataFrames

myfile = "infile.txt"
df = CSV.read(myfile,DataFrame)

```

And see what `df` contains or what warnings come up.

The corresponding starting point in the docs is here: [Home · CSV.jl](https://csv.juliadata.org/stable/index.html#Overview) and starting downwards with the sentence: “That’s quite a bit! Let’s boil down a TL;DR:”

---

<div class="post-metadata">

**Author:** ![ryofurue](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ryofurue/32/24531_2.png) [@ryofurue](https://discourse.julialang.org/u/ryofurue)\
**Post date:** [July 15, 2024, 6:36am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/3 "2024-07-15T06:36:17Z")

</div>

> [@oheil](#):
>
> ```julia-auto
> using CSV, DataFrames
> 
> myfile = "infile.txt"
> df = CSV.read(myfile,DataFrame)
> 
> ```

Thanks! I’ve made some progress using that. But I got stuck with delimiters. According to the thread (from 2021) which I quote at the end of this message, you have to preprocess the input text file if it uses multiple delimiters. Is that still true today? I thought that it was quite usual to expect to be able to specify a Regex, along the lines of

```julia
   df = CSV.read(myfile, DataFrame; delim=r"\s+") # any sequence of "space" characters

```

so that any nonzero sequence of “space” characters ( `\s` ) be tread as one single delimiter.

Currently `CSV.read()` isn’t able to “guess” the number of columns in my datafile, presumably because the datafile uses a mixture of tabs and spaces. `readdlm()` correctly detects the delimiters.

> [@Reading data text files delimited with both spaces & tabs](https://discourse.julialang.org/t/reading-data-text-files-delimited-with-both-spaces-tabs/64851):
>
> Is it possible in CSV.jl, DelimitedFiles.jl, or other, to read text files that contain both spaces & tabs, in the header and/or data sections? An example is provided below. NB: the hidden spaces and tabs should be there after copy and paste, also added row with missing value Col1 Col2 Col3 1 1 012 2 1 013 2 1 2 1 015 Currently, the workaround used is to replace all tabs by spaces in a text editor and then read the file using CSV.jl: using CSV, DataFrames df = CSV.read(file, …

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2024, 8:53am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/4 "2024-07-15T08:53:31Z")

</div>

You are right, this is a bit disappointing and my guess is, the expected convenience is dropped in favor of performance.

At least, consecutive white spaces as delimiting columns should be an option, or in general, consecutive all characters of a given string or a given list of characters would be nice. A regex maybe to much. ~~Perhaps you may open an issue with this request (didn’t check if there is already one).~~ You may support below linked feature request by adding a comment to it.

From your quoted discussion the best solution is doing the editing on the fly using Julia:

```julia
df = read(myfile) |>
        x -> map!(c -> c == UInt8('\t') ? UInt8(' ') : c, x, x) |>
        x -> CSV.read( x, DataFrame; delim=" ", ignorerepeated=true)

```

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2024, 9:33am UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/5 "2024-07-15T09:33:21Z")

</div>

> <https://github.com/JuliaData/CSV.jl/issues/956>
>
> I frequently work with data that uses different delimiters. I would like to be a…ble to store this data as a CSV file without first having to go through and replace the delimiters with just a single choice. The enhancements that I would like to see are for CSV.File as follows:
> 1. The 'delim' argument also accepts an Array of Char or String entries to use as delimiters.
> 2. The 'ignorerepeated' argument could be updated to consider a three state system: 1) All duplicate delimiters ignored 2) Only duplicate delimiters that are the same are ignored and 3) Each delimiter treated as unique
> 
> Thank you for your consideration!

---

<div class="post-metadata">

**Author:** ![ryofurue](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ryofurue/32/24531_2.png) [@ryofurue](https://discourse.julialang.org/u/ryofurue)\
**Post date:** [July 15, 2024, 2:37pm UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/6 "2024-07-15T14:37:36Z")

</div>

By the way, do you know how to tell `CSV.read` to ignore the last delimiter at the end of a line? Without removing it, I get an extra `missing` column.

As a workaround, because I don’t know how to apply two filters to a single stream, I first replace tabs with a space and then remove the line-ending spaces like

```julia
  b = replace(b, r" +\n" => "\n")

```

(which I don’t know will work or not if the text file uses the DOS line ending `\r\n`).

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [July 15, 2024, 3:31pm UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/7 "2024-07-15T15:31:21Z")

</div>

> [@ryofurue](#):
>
> because I don’t know how to apply two filters to a single stream

Perhaps the real problem is, that above solution does not work on strings, but on a `Vector{UInt8}` (array of characters), because that’s what `read` returns into the pipes. Therefor the replacement with `map!`.

To insert filters which work on strings, you can do:

```julia
read(myfile) |> x -> join(Array{Char}(x))

```

This outputs the whole file as a string (probably large) to the REPL.  
For `CSV.read`, we need again the `Vector{UInt8}`:

```julia
collect(UInt8, "line1\nline2")

```

Put it all together:

```julia
df = read(myfile) |>
          x -> map!(c -> c == UInt8('\t') ? UInt8(' ') : c, x, x) |>
          x -> join(Array{Char}(x)) |>
          x -> replace(x, r"\s+\n" => "\n") |>
          x -> collect(UInt8,x) |>
          x -> CSV.read( x, DataFrame; delim=" ", ignorerepeated=true)

```

This issue moved to something quite theoretical by just iterating on a starting problem and sticking to the starting solution and just enhancing it. I am not so happy now with the result and I am not sure if this is still something which performs and scales well.

On the other hand, cleaning up the original source data file manually in an editor is also not a good solution. It’s error prone, can’t be repeated, can’t be undone, it’s intransparent, it’s tedious, … definitely something not recommended by me.

Hopefully, if some other people see this, they may correct me and provide something more appropriate for your kind of _dirty_ data.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 15, 2024, 4:52pm UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/9 "2024-07-15T16:52:17Z")

</div>

> [@ryofurue](#):
>
> This is a very simple interface. You just get a 2D array. How does one use the CSV package to do this?

Note that tabular IO packages in Julia (including CSV) are very flexible in terms of reading into different data structures.  
You don’t generally need something as heavy as DataFrames just to read tables/matrices and work with them:

```julia
using CSV, Tables

CSV.read("file.txt", Tables.matrix) # returns a plain 2d matrix
CSV.read("file.txt", columntable) # returns a simple columnar table: namedtuple of vectors
CSV.read("file.txt", rowtable) # returns a simple row-table: vector of namedtuples

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [July 15, 2024, 6:33pm UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/10 "2024-07-15T18:33:27Z")

</div>

> [@ryofurue](#):
>
> how to tell `CSV.read` to ignore the last delimiter at the end of a line?

Try:  
`stripwhitespace=true`

For example:

```julia
df = CSV.read(IOBuffer(replace(read(inputfile), UInt8('\t') => UInt8(' '))),
     delim=' ', stripwhitespace=true, ignorerepeated=true, comment="#", DataFrame
)

```

---

<div class="post-metadata">

**Author:** ![ryofurue](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ryofurue/32/24531_2.png) [@ryofurue](https://discourse.julialang.org/u/ryofurue)\
**Post date:** [July 17, 2024, 5:45pm UTC](https://discourse.julialang.org/t/reading-text-data-readdlm-is-deprecated-so-how-csv-package-is-used/116955/11 "2024-07-17T17:45:49Z")

</div>

Thank you all for your help!

I confirm that `stripwhitespace=true` removes spaces at the end of the line.

I also confirm that `Tables.matrix` best fits my needs right now. (I just want to access the elements as `df[3, 5]`.)

So, the final solution is

```julia
df = CSV.read( . . . replace \t with ' ' . . . , Tables.matrix; 
  delim=' ', stripwhitespace=true, ignorerepeated=true, comment="#")

```

So, as @oheil says, it would be nice if multiple characters could be specified as `delim`. Then, `CSV.read()` would be as convenient as `readdlm()` for this kind of tabulated, space-delimited text data files.

I did add a comment to the github issue @oheil mentioned.

* * *

[Aside] `CSV.jl`, as its name suggests, is designed for CSV and CSV-like data files, which aren’t optimized for viewing on the computer screen. I often need to use Excel or a similar application just to view a CSV file.

On the other hand, I often encounter tabulated, text data files, which are designed to be intelligible on the command terminal.

I don’t know how those tab characters entered the data file my colleague gave me, but the file is intelligible by just

```bash
$ less thedatafile.txt

```

on the command terminal.
