# Code review: parser and error propagation

**URL:** <https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563>\
**Category:** General Usage\
**Tags:** parsing\
**Created:** [July 28, 2024, 11:01am UTC](https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563 "2024-07-28T11:01:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [July 28, 2024, 11:01am UTC](https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563/1 "2024-07-28T11:01:43Z")

</div>

I need to write parsers for ad-hoc data formats from time to time. I’ve always wondered if I’m propagating errors in a “Julian” way. I don’t know if it’s the best to represent errors by `nothing` as in `Base.tryparse`, a custom error type, or exceptions. Here’s my code for parsing a tuple of two letters separated by a comma. Let me know if the code can be improved.

Tests first, to give an idea about what the code will do:

```julia
@test parse_tuple_of_two_letters("a,b") == ('a', 'b')
@test parse_tuple_of_two_letters("a+b") == ParseError("error in parsing comma: Failed to read character")
@test parse_tuple_of_two_letters("1,b") == ParseError("error in parsing 1st letter: Failed to read character")
@test parse_tuple_of_two_letters("a,2") == ParseError("error in parsing 2nd letter: Failed to read character")
@test parse_letter("a") = 'a'
@test parse_letter("1") = ParseError("Failed to read character")

```

The actual code is below. The code uses a custom error type containing an error string. When the error type is returned by a child function, the parent function adds extra context information to give a sensible error message.

```julia
struct ParseError
    message::String
end

parse_letter(io::IO) = parse_char(io, Base.isletter)
parse_expected_char(io::IO, c::Char) = parse_char(io, ==(c))

function parse_char(io::IO, requirement::Function)::Union{Char, ParseError}
    eof(io) && return ParseError("End of stream!")
    data = read(io, Char)
    if !requirement(data)
        skip(io, -1) # rewind after failed read
        return ParseError("Failed to read character")
    end
    return data
end

parse_tuple_of_two_letters(s::String) = parse_tuple_of_two_letters(IOBuffer(s))

function parse_tuple_of_two_letters(io::IO)::Union{Tuple{Char, Char}, ParseError}
    char1 = parse_letter(io)
    char1 isa ParseError && return ParseError("error in parsing 1st letter: " * char1.message) # return the error in `char1` after context information
    separator = parse_expected_char(io, ',')
    separator isa ParseError && return ParseError("error in parsing comma: " * separator.message) #
    char2 = parse_letter(io)
    char2 isa ParseError && return ParseError("error in parsing 2nd letter: " * char2.message) #
    return (char1, char2)
end

```

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [July 28, 2024, 1:09pm UTC](https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563/2 "2024-07-28T13:09:09Z")

</div>

Make sure to test parsing data with non ascii characters, and invalid unicode.

I would also replace the `read(io, Char)`, and `skip(io, -1)` in `parse_char` with `peek(io, UInt8)` and `skip(io, 1)` if the requirement is met, as not all IO types support negative skipping, and `Char` isn’t always one byte.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [July 28, 2024, 1:40pm UTC](https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563/3 "2024-07-28T13:40:16Z")

</div>

Also, I recommend trying out [GitHub - Seelengrab/Supposition.jl: A Julia implementation of choice sequence based PBT, inspired by Hypothesis](https://github.com/Seelengrab/Supposition.jl) to test potential edge cases you haven’t thought of.

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [July 28, 2024, 3:07pm UTC](https://discourse.julialang.org/t/code-review-parser-and-error-propagation/117563/4 "2024-07-28T15:07:24Z")

</div>

> [@nhz2](#):
>
> not all IO types support negative skipping

Thanks for the reminder. I remember running into this problem when accessing a `.gz` file as a stream decompressed on the fly using `CodecZlib.GzipDecompressorStream`.
