# Is there no standard way to read files with fixed width columns in the new DataFrames ecosystem?

**URL:** <https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324>\
**Category:** Data\
**Created:** [November 26, 2017, 5:39pm UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324 "2017-11-26T17:39:50Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![NickNack](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nicknack/32/1792_2.png) [@NickNack](https://discourse.julialang.org/u/NickNack)\
**Post date:** [November 26, 2017, 5:39pm UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/1 "2017-11-26T17:39:50Z")

</div>

I’ve been using `DataFrames.readtable()` to read files with fixed width columns, but that function is now deprecated in favor of `CSV.read()`. However it seems that `CSV.read()` lacks the flag to “treat consecutive whitespace delimiters as one” that would be required to make it handle fixed width data.

Here’s a [sample file](http://www.pik-potsdam.de/~mmalte/rcps/data/RCP3PD_MIDYEAR_CONCENTRATIONS.DAT) from climate science (standardized scenarios of greenhouse gas concentrations called “Representative Concentration Pathways”). [More similar files available here](http://www.pik-potsdam.de/~mmalte/rcps/index.htm). `DataFrames.readtable()` has no problems with this file:

```julia
readtable("RCP3PD_MIDYEAR_CONCENTRATIONS.DAT"; separator=' ', skipstart=38)

```

But `CSV.read()` fails since it interprets every whitespace as delimiting a new column.

```julia
CSV.read("RCP3PD_MIDYEAR_CONCENTRATIONS.DAT"; delim=' ', header=39)

```

Is there no standard way to read files with fixed width columns in Julia 0.6.1 (or 0.7 or 1.0)? If so that seems to be a pretty serious shortcoming. I know it’s not too hard to write a fixed width file reader from scratch, but I still think this functionality should be included by default. Or am I missing something?

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [November 26, 2017, 10:42pm UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/2 "2017-11-26T22:42:05Z")

</div>

`CSV.jl` hasn’t been updated yet. JuliaData I/O packages like `CSV.jl` and `Feather.jl` are currently being updated.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [November 27, 2017, 6:43am UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/3 "2017-11-27T06:43:01Z")

</div>

> [@NickNack](#):
>
> lacks the flag to “treat consecutive whitespace delimiters as one” that would be required to make it handle fixed width data

Using delimited data tools may be a workaround, but they are not right for the job.

It is very easy to roll your own tools for fixed data, see this topic:

> [@Reading Fixed-Width Column Data](https://discourse.julialang.org/t/reading-fixed-width-column-data/4020):
>
> Hi All, I’m trying to read in a file with rows that look like 13 MEXICO 130971 14 INDONESIA 126582 15 UNITED KINGDOM 114486 Note the possible presence of spaces in the 2nd column. I haven’t been able to find any tool that will easily do this for me (saw [this](https://github.com/JuliaLang/julia/issues/5391), but the suggestion there doesn’t work here). Does one exist that I’m not finding, or do I hav…

That said, fixed width datasets are usually the ugliest to deal with. In the best cases, there is some metadata which tells you the first line of data and the column layout, but there is not widely used standard format (social security data sometimes uses something called `DDF`, but that is not a well-defined format anyway).

I would simply try to get the data in another format, or use a command line tool like `sed` to replace contiguous whitespace with `,` or similar, then read it as `CSV`.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [November 27, 2017, 9:57am UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/4 "2017-11-27T09:57:14Z")

</div>

Indeed, reading fixed-width files as CSV may not be the best approach. We need a dedicated function for that.

---

<div class="post-metadata">

**Author:** ![Mattriks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattriks/32/351_2.png) [@Mattriks](https://discourse.julialang.org/u/Mattriks)\
**Post date:** [November 27, 2017, 12:02pm UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/5 "2017-11-27T12:02:56Z")

</div>

Here is quick workaround (probably not a general solution), but it works for me (Julia 0.6.1) using the sample file above:

```nohighlight
function readfwf(x::String; opts...) 
    d = readdlm(x; opts...)
    z = !isa(d,Tuple) ? DataFrame(d) : names!(DataFrame(d[1]), Symbol.(vec(d[2])))
    return z
end    

x = "RCP3PD_MIDYEAR_CONCENTRATIONS.DAT"
d = readfwf(x, header=true, skipstart=38)

```

---

<div class="post-metadata">

**Author:** ![NickNack](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nicknack/32/1792_2.png) [@NickNack](https://discourse.julialang.org/u/NickNack)\
**Post date:** [November 28, 2017, 2:15pm UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/6 "2017-11-28T14:15:40Z")

</div>

Thanks for the suggestions guys, but I had already solved my problem when I posted (not as elegantly as @Mattriks’ little function though). I was just surprised that some of the functionality of `readtable()` had fallen through the cracks during the transition to the new DataFrames version and wanted to shine some light on it.

---

<div class="post-metadata">

**Author:** ![Diego\_Javier\_Zea](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/diego_javier_zea/32/1858_2.png) [@Diego\_Javier\_Zea](https://discourse.julialang.org/u/Diego_Javier_Zea)\
**Post date:** [November 29, 2017, 11:32am UTC](https://discourse.julialang.org/t/is-there-no-standard-way-to-read-files-with-fixed-width-columns-in-the-new-dataframes-ecosystem/7324/7 "2017-11-29T11:32:25Z")

</div>

Parsing fixed width columns depending on a variable number of spaces could create (sometimes silent) bugs when there is no space between fields. In the best scenario, this will end up in an error (e.g., an error for having a shorter row).
