# Stricter date parsing

**URL:** <https://discourse.julialang.org/t/stricter-date-parsing/12334>\
**Category:** General Usage\
**Tags:** question, dates, parsing\
**Created:** [July 12, 2018, 2:40pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334 "2018-07-12T14:40:27Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 12, 2018, 2:40pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/1 "2018-07-12T14:40:27Z")

</div>

I don’t know if this is intended or not, but date parsing is somewhat too permissive:

```julia
julia> VERSION
v"0.7.0-beta.275"

julia> using Dates

julia> tryparse(Date, "19800101", dateformat"yyyy-mm-dd")
19800101-01-01

```

Is this a bug or a feature? How can I get stricter validation?

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [July 12, 2018, 3:30pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/2 "2018-07-12T15:30:39Z")

</div>

Yikes, that seems rather iffy! Do date formats implicitly not require the entire format to match?

---

<div class="post-metadata">

**Author:** ![rened](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rened/32/19968_2.png) [@rened](https://discourse.julialang.org/u/rened)\
**Post date:** [July 12, 2018, 4:31pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/3 "2018-07-12T16:31:39Z")

</div>

It might not be as clear cut - what about 123 A.D.?

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [July 12, 2018, 4:42pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/4 "2018-07-12T16:42:45Z")

</div>

I would definitely call that a bug, you should probably file an issue.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [July 12, 2018, 5:04pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/5 "2018-07-12T17:04:39Z")

</div>

The bigger issue to me is that it ignores the trailing `-mm-dd` part of the date format. That part of the data seems like it would be mandatory, not optional. The year could potentially be 19800101, although it might be sensible to only parse years as long as the given number of format digits.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 12, 2018, 8:38pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/6 "2018-07-12T20:38:25Z")

</div>

Thanks for the discussion. Opened issue  
[https://github.com/JuliaLang/julia/issues/28090](https://github.com/JuliaLang/julia/issues/28090)

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [July 12, 2018, 8:43pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/7 "2018-07-12T20:43:43Z")

</div>

The DateFormat docs could probably be clearer here, but essentialy `dateformat"y-m-d"` ends up parsing the exact same dates as `dateformat"yyyy-mm-dd"`; the presence of a `y` or `m` just means, “parse the digits until another delimiter is encountered”. So the doc clarification should mention there’s currently no way to _restrict_ the number of digits parsed.

There’s also not currently a way to mark a date part as “mandatory”; if some parts are parsed, then defaults are assumed for the rest. That just goes back to the fact that we have `Date(2018, 1)`, which defaults to January 1st, 2018, even though no “day” part was given.

So while this might be surprising and maybe some enhancements need to be made, there’s not really any “bugs” here; this is definitely working as intended.

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [July 12, 2018, 8:49pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/8 "2018-07-12T20:49:34Z")

</div>

> [@quinnj](#):
>
> So while this might be surprising and maybe some enhancements need to be made, there’s not really any “bugs” here; this is definitely working as intended.

Makes sense to me.

I take it this is standard date parsing behavior and others do this as well?

This definitely seems like it would be a nasty surprise for somebody somewhere, but if it is some sort of universal standard behavior for date parsing the burden would seem to be on user. If it is a Julia specific quirk the burden would seem to be on Julia.

(I was going to check what Python does but I can’t be bothered looking up how to do it right now lol)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 12, 2018, 8:49pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/9 "2018-07-12T20:49:46Z")

</div>

> [@quinnj](#):
>
> there’s not really any “bugs” here; this is definitely working as intended

Possibly, but at least the documentation should be improved then. `?DateFormat` says

> `yyyymmdd 19960101 Matches fixed-width year, month, and day`

which may lead the user to conclude that repeating characters leads to fixed-width parsing.

Also, it can be argued that if the parser is not validating things like `yyyy-mm-dd`, perhaps it should, or at least syntax should be implemented for it. But that’s a wishlist item.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [July 12, 2018, 8:56pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/10 "2018-07-12T20:56:17Z")

</div>

Yeah, there’s just a big difference parsing-wise between fixed-width and delimited. With fixed-width, obviously you’re expecting only a certain # of characters and more or less than what is specified is going to blow things up.

With delimited though, what if you had a mix of date strings like `["2018-1-1", "2018-1-2", "2018-1-20"]`, how do you specify the “day” part for parsing? You either of a dateformat string like `"yyyy-mm-dd"` or `"yyyy-m-d"`, but you can’t really _require_ one or two digits, because there might be both in a set of date strings to be parsed.

So, generally, the approach is that fixed-width parsing is always more strict, while delimited parsing is more lenient.

Another thing maybe we should consider is taking away the default values for month & day parts; that would mean you’d have to explicitly call `Date(2018, 1, 1)` instead of relying on defaults. That would have helped this situation because `Date(19800101)` would have thrown an error because no month or day arguments were given. I’m not sure how widely used/expected/relied upon those default arguments are though.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 13, 2018, 8:46am UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/11 "2018-07-13T08:46:02Z")

</div>

> [@quinnj](#):
>
> Another thing maybe we should consider is taking away the default values for month & day parts; that would mean you’d have to explicitly call `Date(2018, 1, 1)` instead of relying on defaults. That would have helped this situation because `Date(19800101)` would have thrown an error because no month or day arguments were given. I’m not sure how widely used/expected/relied upon those default arguments are though.

I would be OK with keeping the default options in the `Date` constructor, but _not_ have the parser fill in missing fields. When I parse a `dateformat"y-m-d"`, I implicitly expect an `m` and `d` to be there, also the two `-`s.

A few words on my actual use case: sometimes I parse datasets from CSV or similar where someone thought that having sentinel values like `99999` is a great idea… even in non-numerical columns. I don’t want this to parse as `Date(99999)`.

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [July 13, 2018, 1:32pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/12 "2018-07-13T13:32:06Z")

</div>

> [@quinnj](#):
>
> how do you specify the “day” part for parsing? You either of a dateformat string like `"yyyy-mm-dd"` or `"yyyy-m-d"` , but you can’t really _require_ one or two digits, because there might be both in a set of date strings to be parsed.

I think what I get hung up on is that since the characters get repeated (i.e. `"yyyy"`) I expect that to indicate a specific width or at least an upper limit, though I do vaguely seem to remember there being different characters for fixed width.

> [@Tamas\_Papp](#):
>
> A few words on my actual use case: sometimes I parse datasets from CSV or similar where someone thought that having sentinel values like `99999` is a great idea… even in non-numerical columns. I don’t want this to parse as `Date(99999)` .

I usually find that it’s better to replace those values with `missing` (or whatever) before trying to parse the set.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 13, 2018, 1:34pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/13 "2018-07-13T13:34:53Z")

</div>

> [@ExpandingMan](#):
>
> I usually find that it’s better to replace those values with `missing` (or whatever) before trying to parse the set.

Eventually, yes, but first I have to know what they are. Imagine a half-TB dataset with undocumented conventions for missing values that change occasionally. For the first pass, I just want to know if something is a date of a certain kind, or not, collect the invalid values and look at them.

In any case, I think that restricting the parsing to be less permissive in the direction I suggested is generally useful.

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [July 13, 2018, 1:37pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/14 "2018-07-13T13:37:48Z")

</div>

> [@Tamas\_Papp](#):
>
> Eventually, yes, but first I have to know what they are. Imagine a half-TB dataset with undocumented conventions for missing values that change occasionally. For the first pass, I just want to know if something is a date of a certain kind, or not, collect the invalid values and look at them.

I feel your pain. This is definitely a case for very strict parsing.

By the way, do we have any methods that return `missing` in the event of a parsing error? `try` `catch` can be rather slow, as I understand.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 13, 2018, 1:43pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/15 "2018-07-13T13:43:30Z")

</div>

`tryparse` returns `nothing` (in v0.7), which works fine. It is very fast.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [July 13, 2018, 2:19pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/16 "2018-07-13T14:19:03Z")

</div>

Do you need it just for this one date format?

```nohighlight
julia> checkdate(date::String) = ismatch(r"[0-9]{1,4}-[0-9]{1,2}-[0-9]{1,2}",date)
checkdate (generic function with 1 method)

julia> checkdate("2018-07-13")
true

julia> checkdate("19800101")
false

```

Or do you need something that works for a generic date format specification?

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [July 13, 2018, 2:29pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/17 "2018-07-13T14:29:49Z")

</div>

Here is a slightly more generic solution

```nohighlight
julia> function checkdate(date::String,format::String)
           l = length.(split(format,'-'))
           r = join(["[0-9]{1,$(l[k])}" for k ∈ 1:length(l)],'-')
           ismatch(Regex("^$r\$"),date)
       end;

julia> checkdate("2018-07-13","yyyy-mm-dd")
true

julia> checkdate("19800101","yyyy-mm-dd")
false

```

This actually sets an upper limit to the number of digits after counting them.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 13, 2018, 2:39pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/18 "2018-07-13T14:39:31Z")

</div>

I would prefer something generic. Also note that since `v0.6`, date parsing using `DateFormat` is very heavily optimized; I did not benchmark but I imagine a regex-based solution would be orders of magnitude slower.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [July 13, 2018, 2:44pm UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/19 "2018-07-13T14:44:16Z")

</div>

Agree, I don’t see why constructor defaults and allowing incomplete parsing are coupled. It seems reasonable for constructors to supply defaults while format parsers require the entire format to be matched (and anything not in the format to be supplied by default). The current behavior strikes me as dangerous enough to be considered something of a design bug.

---

<div class="post-metadata">

**Author:** ![Liso](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@Liso](https://discourse.julialang.org/u/Liso)\
**Post date:** [July 14, 2018, 7:24am UTC](https://discourse.julialang.org/t/stricter-date-parsing/12334/20 "2018-07-14T07:24:13Z")

</div>

This could be probably interesting here too:

> [@ANN: Automa.jl - a package to compile regular expressions to Julia](https://discourse.julialang.org/t/ann-automa-jl-a-package-to-compile-regular-expressions-to-julia/1586):
>
> Hi all, I released a new package, [Automa.jl](https://github.com/BioJulia/Automa.jl), which is a set of tools to compile regular expressions written in Julia DSL into optimized Julia code. Here is a short example to tokenize a string into numerical literals from [example/numbers.jl](https://github.com/BioJulia/Automa.jl/blob/master/example/numbers.jl): import Automa import Automa.RegExp: @re\_str const re = Automa.RegExp # Describe patterns in regular expression. oct = re"0o[0-7]+" dec = re"[-+]?[0-9]+" hex = re"0x[0-9A-Fa-f]+" prefloat = re"[-+]?([0-9]+\.[0-9]\*|[0-9]\*\.[0-9]+)" float =…

[Next page](https://discourse.julialang.org/t/stricter-date-parsing/12334.md?page=2)
