# (The future of) HTML Parsing in Julia

**URL:** <https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337>\
**Category:** Web Stack\
**Tags:** question, package\
**Created:** [June 14, 2023, 3:02pm UTC](https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337 "2023-06-14T15:02:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 14, 2023, 3:02pm UTC](https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337/1 "2023-06-14T15:02:56Z")

</div>

Hello world,

I am writing this here hoping to find someone more knowledgeable than me on the Julia ecosystem around web / HTML.

My context in a few words: I am working on a larger project that involves lots of HTML manipulations (reliable parsing is paramount). I started the project by using the [Gumbo.jl](https://github.com/JuliaWeb/Gumbo.jl) package, and I realized a while back that the C counterpart was archived. Since Gumbo.jl is a wrapper around the C version, I think it is safe to assume that it will stay frozen and not implement new features dealing with new HTML standard changes/additions. Continuing to build on Gumbo.jl seems kind of risky at this point.

I was also looking at [EzXML.jl](https://github.com/JuliaIO/EzXML.jl) package - but that one also seems not maintained (some community pull requests are just standing there without being merged).

Don’t get me wrong: I do not want people to serve me “the parser” on a plate - I want to understand the direction of web programming in Julia - and an HTML parser seems to me like a kind of cornerstone. I am also willing to contribute, but I don’t want to bet on the wrong horse here: maybe there is a direction that has a better chance of succeeding in the long run, and having just another small Julia package without documentation or contributors is not going to cut it.

Maybe some of the maintainers of Genie.jl can help with this. I understand that Genie.jl has some internals that deals with HTML parsing - but it seems that the functionality is tailored to templating and MVC approach. Maybe a solution to have a standalone HTML parser is to separate the parsing from Genie.jl as a standalone package.

However, can anybody help with some directions here?

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [June 14, 2023, 3:10pm UTC](https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337/2 "2023-06-14T15:10:01Z")

</div>

There are a few other packages that you didn’t list that could be helpful.

> **[GitHub - JuliaComputing/XML.jl: Read and write XML in pure Julia](https://github.com/JuliaComputing/XML.jl)**
>
> Read and write XML in pure Julia. Contribute to JuliaComputing/XML.jl development by creating an account on GitHub.

> **[GitHub - JuliaPluto/HypertextLiteral.jl: Julia library for the string...](https://github.com/JuliaPluto/HypertextLiteral.jl)**
>
> Julia library for the string interpolation of HTML and SVG - GitHub - JuliaPluto/HypertextLiteral.jl: Julia library for the string interpolation of HTML and SVG

[GitHub - tlienart/Franklin.jl: (yet another) static site generator. Simple, customisable, fast, maths with KaTeX, code evaluation, optional pre-rendering, in Julia.](https://github.com/tlienart/Franklin.jl) Also must have some capabilities in this area I’m guessing.

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 15, 2023, 3:10am UTC](https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337/3 "2023-06-15T03:10:14Z")

</div>

Thanks for pointing out the additional materials.

I was looking specifically for the HTML parsing functionality - although HypertextLiteral.jl looks really great for generating HTML on the fly (and I wasn’t aware of the package).

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 15, 2023, 3:41am UTC](https://discourse.julialang.org/t/the-future-of-html-parsing-in-julia/100337/4 "2023-06-15T03:41:07Z")

</div>

I was aware of the [Cobweb.jl](https://github.com/JuliaComputing/Cobweb.jl) - but I actually needed to dive into the source code to find the parser (the documentation doesn’t say anything about that).

I just tested it - it seems a feasible and easy-to-extend/maintain alternative to Gumbo.jl project.
