# Reading HTML file for parsing

**URL:** <https://discourse.julialang.org/t/reading-html-file-for-parsing/91873>\
**Category:** General Usage\
**Created:** [December 19, 2022, 7:41pm UTC](https://discourse.julialang.org/t/reading-html-file-for-parsing/91873 "2022-12-19T19:41:58Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![edas](https://avatars.discourse-cdn.com/v4/letter/e/958977/32.png) [@edas](https://discourse.julialang.org/u/edas)\
**Post date:** [December 19, 2022, 7:41pm UTC](https://discourse.julialang.org/t/reading-html-file-for-parsing/91873/1 "2022-12-19T19:41:59Z")

</div>

I want to read html file and then parse it (not sure if use word “parse” correctly).

I saved [example.com](http://example.com) to file example.html  
Also using EzXML library

```julia
using Cascadia, Gumbo, HTTP,AbstractTrees
using EzXML

r = EzXML.readhtml("example.html")

print(r)

```

Prints result (html):

\<?xml version="1.0" encoding="UTF-8" standalone="yes"?\> Example Domain

```
<meta charset="utf-8"/>
<meta http-equiv="Content-type" content="text/html; charset=utf-8"/> <meta name="viewport" content="width=device-width, initial-scale=1"/>
<style type="text/css"><![CDATA[
body {
    background-color: #f0f0f2;
    margin: 0;
    padding: 0;
    font-family: -apple-system, system-ui, BlinkMacSystemFont, "Segoe UI", "Open Sans", "Helvetica Neue", Helvetica, Arial, sans-serif;     

}
div {
    width: 600px;
    margin: 5em auto;
    padding: 2em;
    background-color: #fdfdff;
    border-radius: 0.5em;
    box-shadow: 2px 3px 7px 2px rgba(0,0,0,0.02);
}
a:link, a:visited {
    color: #38488f;
    text-decoration: none;
}
@media (max-width: 700px) {
    div {
        margin: 0 auto;
        width: auto;
    }
}
]]></style>

```

# Example Domain

This domain is for use in illustrative examples in documents. You may use this domain in literature without prior coordination or asking for permission.

[More information...](https://www.iana.org/domains/example)

How to work further with this html ?

```julia
h = parsehtml(String(r.body))

```

Gives error:  
ERROR: type Document has no field body

Cascadia commands does not work too.

EzXML.readhtml() reads simple html files, but give errors on more complex files.

What library should I use ?  
Or I missed some steps ?

Thanks

---

<div class="post-metadata">

**Author:** ![avik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/avik/32/17_2.png) [@avik](https://discourse.julialang.org/u/avik)\
**Post date:** [December 19, 2022, 8:10pm UTC](https://discourse.julialang.org/t/reading-html-file-for-parsing/91873/2 "2022-12-19T20:10:20Z")

</div>

You seem to be importing [Gumbo.jl](https://juliahub.com/ui/Packages/Gumbo/mllB2/0.8.2) and [Cascadia.jl](https://juliahub.com/ui/Packages/Cascadia/Pq6Fi/1.0.2) in your code. So I don’t really understand why you are using EzXML’s parser? Can’t you directly use Gumbo’s [`parsehtml` method](https://juliahub.com/ui/Search?q=parsehtml&type=symbols&u=define&p=Gumbo%3A708ec375-b3d6-5a57-a7ce-8257bf98657a)?

XML based parsers such as EzXML often have difficulty working with actual html.
