# More control over writedlm formatting

**URL:** <https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840>\
**Category:** Internals & Design\
**Tags:** proposal\
**Created:** [March 23, 2017, 5:45pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840 "2017-03-23T17:45:37Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![mcopik](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@mcopik](https://discourse.julialang.org/u/mcopik)\
**Post date:** [March 23, 2017, 5:45pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/1 "2017-03-23T17:45:37Z")

</div>

Hi,

I’ve recently discovered that writedlm() tends to output floating-point numbers using a very strange format pattern which removes the fractional part by decreasing exponent. Not only it is terribly confusing, the data is no longer easy to read and understand by a human.

An example:  
`6.0444e-5 -> 60444e-9 2.15234e-5 -> 215234e-10 3.3253e-5 -> 33253e-9`  
It is not a bug - `print_shortest` has done it job by reducing the output by exactly one character - a decimal mark. But it is not human readable anymore and it can take a lot of time to notice that decimal mark is gone and the exponent is indeed correct.

I have seen that a similar problem [with print function](https://github.com/JuliaLang/julia/issues/6452) has been resolved by not using `print_shortest` anymore and it looks that I’m neither [first](http://julia-programming-language.2336112.n4.nabble.com/Specify-format-in-writedlm-td40590.html) nor [second](https://groups.google.com/forum/#!topic/julia-users/YsO2g2HcL5c) to ask questions due to lack of configurability in this function. [The source code](https://github.com/JuliaLang/julia/blob/59b6080e0a75dbe88b4890189ef02f2dbf4164ea/base/datafmt.jl#L488) suggests that at some point (at least two years ago) there was an idea to extend capabilities of writedlm. I have to implement a new functionality because the current situation is sadly not acceptable for us and I might just try to contribute to Julia and extend writedlm(). The questions here are: is there some solution which I have not been able to find? Is there someone already working on that? Do you already have an idea how to extend `writedlm` signature or do you have plans for a more generic IO formatting?

Another option here is to modify `print_shortest` but I guess that there is a reason why it is being used.

Best regards,  
Marcin

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [March 24, 2017, 1:31pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/2 "2017-03-24T13:31:15Z")

</div>

Use the CSV package instead: [Home · CSV.jl](http://juliadata.github.io/CSV.jl/latest/). We should really delete readdlm and writedlm.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [March 24, 2017, 2:12pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/3 "2017-03-24T14:12:32Z")

</div>

> [@StefanKarpinski](#):
>
> We should really delete readdlm and writedlm.

Please, don’t. Probably CSV can do lots more than `readdlm`, but the latter can do easily very simple thing, without the need of using a `DataFrame`. `CSV.read` is not even able to read an UTF-encoded file correctly[1], I have no problem with `readdlm`.

I remember you also [expressed similar “disdain”](https://github.com/JuliaLang/julia/issues/10649) for QuadGK, which is actually a very nice piece of code 😉

Note:  
[1] This seems specific to CSV.jl, DataFrames.jl works flawless.

Edit: the problem with `CSV.jl` can be actually solved by setting the keyword `weakrefstrings=false`.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 24, 2017, 3:10pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/4 "2017-03-24T15:10:02Z")

</div>

Isn’t `readdlm` for reading `Matrix`es directly? I think that functionality is distinct from CSV’s? (I don’t think a `Matrix` defines a Data.Sink)

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [March 24, 2017, 8:35pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/5 "2017-03-24T20:35:51Z")

</div>

It has nothing to do with disdain. Numerical integration just doesn’t belong in a language’s standard library. It’s fine to include it in a standard distribution of packages, but there’s just no reason to have integration in the base language. In the case of CSV reading and writing, it’s for a different reason: the current state of affairs is perfectly demonstrated by this thread. People naturally try the built-in functionality which lags behind the CSV package in both features and performance, and whenever anyone has a question the answer ends up being to use the CSV package. The fact that CSV.read doesn’t handle UTF-8 data incorrectly doesn’t mean that we should continue to split efforts between readcsv and CSV – it means that CSV should be fixed and readcsv should be deleted.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [March 24, 2017, 8:38pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/6 "2017-03-24T20:38:27Z")

</div>

In general, we should not have several half-baked solutions to problems – be it integration, csv reading/writing, or whatever. Instead, we should strive to have a single good solution. Having a half-baked solution in the standard library actively prevents us from reaching that better state of affairs.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [March 24, 2017, 8:54pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/7 "2017-03-24T20:54:21Z")

</div>

“Disdain” was between quotes for a reason 😉

Anyway, I see your point, but really, for many simple uses `readdlm` and `writedlm` are very handy. Of course, for more fine-grained control other tools are more suited.

How about moving to external package at least? They wouldn’t look as “the official way for reading/writing matrices in Julia”, but at least would still be available.

---

<div class="post-metadata">

**Author:** ![Paul\_Soderlind](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paul_soderlind/32/1753_2.png) [@Paul\_Soderlind](https://discourse.julialang.org/u/Paul_Soderlind)\
**Post date:** [March 24, 2017, 9:15pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/8 "2017-03-24T21:15:58Z")

</div>

It is probably reasonable to move readdlm and similar things out of the standard library, but I would kindly ask you to offload them to simple and stand-alone packages. QuadGK is a good example.

There is something to be said for packages that do not have lots of cross-dependencies (on other packages).

/Paul S

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 24, 2017, 9:47pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/9 "2017-03-24T21:47:22Z")

</div>

I think there might be some confusion about words here? I am sure almost everybody agrees that `CSV.read` is superior to `readcsv` and `readtable` for reading DataFrames, and that these should be removed in favour of a CSV dependency. But the question was about `readdlm` / `writedlm` which are also used for directly reading numerical matrices into the `Matrix` type. That is still really useful.

Or am I misunderstanding this? (EDIT I was misrepresenting this slightly, so I edited the text above).

---

<div class="post-metadata">

**Author:** ![mzaffalon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mzaffalon/32/214168_2.png) [@mzaffalon](https://discourse.julialang.org/u/mzaffalon)\
**Post date:** [March 25, 2017, 4:42am UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/10 "2017-03-25T04:42:07Z")

</div>

If this happens, are there plans to add a section in the official documentation with a collection of recommended libraries as opposed to listing 1300+ packages and their build status? For somebody not keeping a close eye on the current state of development, an extended standard library documentation section would be the fastest way to get started.

(Since it was mentioned, `quadgk` is not a standard library routine any longer in v0.6 and the only reference I could find to the Quadgk.jl package is in the developer section about compiler efficiency.)

---

<div class="post-metadata">

**Author:** ![mcopik](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@mcopik](https://discourse.julialang.org/u/mcopik)\
**Post date:** [March 25, 2017, 3:34pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/12 "2017-03-25T15:34:16Z")

</div>

> [@mkborregaard](#):
>
> But the question was about readdlm / writedlm which are also used for directly reading numerical matrices into the Matrix type. That is still really useful.

I agree fully. I believe it is a very common problem and having a simple file I/O for numerical matrices would be very beneficial to users.

The real question here is not if writedlm should be more powerful or should be abandoned. I’m still not understanding why `215234e-10` is considered to be a better format than `2.15234e-5`. It’s not easier to read, it’s barely shorter and it is terribly confusing. Is there anyone else here who agrees with me that this is just a bug in `print_shortest`?

---

<div class="post-metadata">

**Author:** ![mcopik](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@mcopik](https://discourse.julialang.org/u/mcopik)\
**Post date:** [March 25, 2017, 3:57pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/13 "2017-03-25T15:57:03Z")

</div>

> [@StefanKarpinski](#):
>
> People naturally try the built-in functionality which lags behind the CSV package in both features and performance, and whenever anyone has a question the answer ends up being to use the CSV package.

> [@StefanKarpinski](#):
>
> In general, we should not have several half-baked solutions to problems – be it integration, csv reading/writing, or whatever. Instead, we should strive to have a single good solution. Having a half-baked solution in the standard library actively prevents us from reaching that better state of affairs.

I understand your POV but, as I said above in another reply, saving a small array of numbers to file is a very common task which does not require a huge flexibility or high throughput and using a full library there is an overkill. One can argue that my complaint is based on a lack of configurability in writedlm() but I did not expect from it to be flexible or efficient, I expected a simple utility to save data in a human readable format and it failed me.

Do you think that I should download an additional and much more complex package to just create a datafile for pgfplots? I’d have to put this burden on each user of my very small and simple library. I downloaded the package and here is the output:

```julia
INFO: Installing CSV v0.1.2
INFO: Installing CategoricalArrays v0.1.3
INFO: Installing Compat v0.21.0
INFO: Installing DataArrays v0.4.0
INFO: Installing DataFrames v0.9.0
INFO: Installing DataStreams v0.1.2
INFO: Installing DataStructures v0.5.3
INFO: Installing FileIO v0.3.1
INFO: Installing GZip v0.3.0
INFO: Installing NullableArrays v0.1.0
INFO: Installing Reexport v0.0.3
INFO: Installing SortingAlgorithms v0.1.1
INFO: Installing SpecialFunctions v0.1.1
INFO: Installing StatsBase v0.13.1
INFO: Installing WeakRefStrings v0.2.0

```

Fifteen packages to create simple text file with few strings and numbers, and it is very likely that I will never use any other functionality from these packages in my project. I can achieve the same thing with one function call in NumPy (savetxt) or two in MATLAB (fprintf for header and dlmwrite).

The most likely scenario here is that I’ll just implement a loop over IOBuffer generating a very long string representing my data. It won’t be efficient but it doesn’t have to. It won’t be elegant but it will keep my library simple and easy to use.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 25, 2017, 6:37pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/14 "2017-03-25T18:37:58Z")

</div>

What’s the actual problem with installing these 15 packages? They are all pure Julia and very lightweight. They don’t make your library complex nor hard to use. At some point there will be a set of preinstalled standard packages, in which CSV might be included.

The problem with having simple default functions is that we keep having to tell people to use the more complex implementation when the basic one doesn’t suit their needs. This wastes everybody’s time.

It should be possible to allow CSV.jl to return data as a matrix for when you don’t want to use a data frame. The code for that just needs to be written.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 25, 2017, 7:00pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/15 "2017-03-25T19:00:54Z")

</div>

Though I may not care about installing 15 packages (especially basic ones like these) it still seems like a big dependency to incur for a package to be able to read a CSV file, given the efforts we otherwise take to keep packages dependency-light and modular (e.g. StatPlots is a distinct package from Plots precisely so that Plots can avoid depending on those packages).

My main concern was to point out that `readcsv` and `CSV.read` do not have completely overlapping functionalities, which should be considered if `readcsv` should be deleted. That would be remedied if the `Matrix` type would define a `Data.Sink` as you say. But is that really the best design?

Would it not seem more obvious to make the methods distinct by deprecating `readcsv`s ability to read DataFrames and other general table structures (via the ::Type positional argument), but keep the function as a way to read and write a `Matrix` in Base; and then making CSV.read the only way of reading DataFrames (and friends)? It would still save the issue of having to tell people to use a different function.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 25, 2017, 7:11pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/16 "2017-03-25T19:11:24Z")

</div>

(maybe the discussion about deprecating `readdlm` should be split into a different topic and this thread kept to discuss the issue of `print_shortest` formatting)

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [March 25, 2017, 8:00pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/17 "2017-03-25T20:00:53Z")

</div>

Having functionality for loading an array of delimited values with a single common element type, with no support for escaping data or anything complex. seems fine, but `readddlm` and `writedlm` try to do way too much more than that. The number of dependencies of CSV.jl is an issue, but the fact that it _currently_ has too many dependencies is not really an argument that we shouldn’t strive to have a simpler, better factored, generic CSV reading/writing in an external package instead of in the base library.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 25, 2017, 8:26pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/18 "2017-03-25T20:26:27Z")

</div>

I agree completely with that.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [March 25, 2017, 9:23pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/19 "2017-03-25T21:23:46Z")

</div>

The dependencies situation for CSV is indeed currently unfortunate, but it’s entirely due to the current dependency on DataFrames, which has dependency-bloated over the years and is the number one example I’m aware of that would benefit from optional dependencies. I could remove the DataFrames dependency, which means users would have to write `CSV.read(file, DataFrame)` instead, which isn’t terrible.

Also, there’s absolutely no reason we can’t define `Data.Sink` interface methods for `Matrix` and it’s been on my list for a while, just haven’t gotten around to.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [March 25, 2017, 9:30pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/20 "2017-03-25T21:30:52Z")

</div>

Most of the dependencies seem to also come through DataStreams though? Or could you also remove DataFrames and the others from that package? (don’t be mistaken, I am a huge fan of CSV.jl 🙂 )

---

<div class="post-metadata">

**Author:** ![mcopik](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@mcopik](https://discourse.julialang.org/u/mcopik)\
**Post date:** [March 27, 2017, 12:48pm UTC](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840/21 "2017-03-27T12:48:09Z")

</div>

> [@mkborregaard](#):
>
> (maybe the discussion about deprecating readdlm should be split into a different topic and this thread kept to discuss the issue of print\_shortest formatting)

Do you think I should start another thread here or open an issue on Github?

[Next page](https://discourse.julialang.org/t/more-control-over-writedlm-formatting/2840.md?page=2)
