# Reading fixed-width files?

**URL:** <https://discourse.julialang.org/t/reading-fixed-width-files/30163>\
**Category:** General Usage\
**Created:** [October 22, 2019, 11:16am UTC](https://discourse.julialang.org/t/reading-fixed-width-files/30163 "2019-10-22T11:16:37Z")\
**Posts on this page:** 3\
**Page:** 2

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [March 22, 2020, 12:18pm UTC](https://discourse.julialang.org/t/reading-fixed-width-files/30163/21 "2020-03-22T12:18:49Z")

</div>

> [@bkamins](#):
>
> noted that in practice problems with multi-byte characters could be encountered

I may misunderstand the issue you linked, but I don’t see the example where this happens _in practice_. It just raises the possibility.

Some formats, eg recent incarnations of Stata’s `dta` format, store UTF8 in fixed width fields. The way that works is that UTF8 is just considered a byte string, which is then padded/read as is. Eg `"ηβπ"` would take 6 bytes.

This is pretty much the only format that makes sense. “Fixed width” in _characters_ coded in a variable-length encoding like UTF8 (which is pretty much all that should be practically relevant, even though it is easy to support just about anything else in Julia) makes no sense as it throws out all the actual _advantages_ of fixed width.

Again, I am sure there is _someone_ using that to store data. But it is not something a sane library would even consider supporting because it requires an entirely different approach.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [March 22, 2020, 12:22pm UTC](https://discourse.julialang.org/t/reading-fixed-width-files/30163/22 "2020-03-22T12:22:15Z")

</div>

While all the widths should (probably?) indeed be fixed in bytes, from a user POV it makes sense to _specify_ these widths in terms of characters. I.e. if I open the table file in a text editor, I can only see and count characters - not bytes. And these character counts are what should be specified as column widths and positions.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 22, 2020, 1:35pm UTC](https://discourse.julialang.org/t/reading-fixed-width-files/30163/23 "2020-03-22T13:35:19Z")

</div>

The issue I have linked is just a summary of the discussion where if I recall correctly such files were occuring in practice (if I am not mistaken they were generated using COBOL on mainframes).

Anyway FWF.jl by default used byte width but it can be switched using a kwarg. I guess the simplest thing to do for someone interested in having a common FWF reader/writer is to make a PR to [GitHub - RandomString123/FWF.jl: Fixed width file parsing in Julia](https://github.com/RandomString123/FWF.jl) to make it work on modern Julia.

In the long run probably having it in CSV.jl, if @quinnj would consider this, would be the best option as there is loads of parsing functionality already in CSV.jl that is vastly superior to FWF.jl.

[Previous page](https://discourse.julialang.org/t/reading-fixed-width-files/30163.md?page=1)
