# charset=Windows-1250, Gumbo, parsing html , Hoow to keep orginal text

**URL:** <https://discourse.julialang.org/t/charset-windows-1250-gumbo-parsing-html-hoow-to-keep-orginal-text/9084>\
**Category:** General Usage\
**Created:** [February 15, 2018, 4:23pm UTC](https://discourse.julialang.org/t/charset-windows-1250-gumbo-parsing-html-hoow-to-keep-orginal-text/9084 "2018-02-15T16:23:43Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [February 15, 2018, 4:23pm UTC](https://discourse.julialang.org/t/charset-windows-1250-gumbo-parsing-html-hoow-to-keep-orginal-text/9084/1 "2018-02-15T16:23:43Z")

</div>

I am using this code to parsing html pages:  
’  
using Gumbo  
using AbstractTrees  
url=“[http://rp.pl](http://rp.pl)”  
getpage(url) = parsehtml(String(read(download(url))))  
text\_only(doc::HTMLDocument) = text\_only(doc.root)  
text\_only(frag) = join([text(leaf) for leaf in Leaves(frag) if leaf isa HTMLText], " ")  
get\_page\_text(url) = text\_only(getpage(url))  
doc=parsehtml(String(read(download(url))));  
only=sort(split(text\_only(doc.root[2]) ))  
’  
If charset=Windows-1250 for this code all natiopnal Char are lost. E.g. życie =\> �ycie etc.

What teke pages with charset=Windows-1250 ? e.g url=“[http://rp.pl](http://rp.pl)”  
Paul

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 15, 2018, 6:04pm UTC](https://discourse.julialang.org/t/charset-windows-1250-gumbo-parsing-html-hoow-to-keep-orginal-text/9084/2 "2018-02-15T18:04:12Z")

</div>

Julia’s `String` type expects UTF-8 encoded text. Non-ASCII Windows-1252 text will show up as [mojibake](https://en.wikipedia.org/wiki/Mojibake) or invalid UTF-8.

If you have Windows-1250 text (or other encodings) then you have a few options:

- Convert to UTF-8. (A package like [https://github.com/nalimilan/StringEncodings.jl](https://github.com/nalimilan/StringEncodings.jl) can help here, although for the specific case of Windows-1250 you can probably make something more efficient if needed.)
- Use `String` on the Windows-1252 text as-is. Non-ASCII characters will not be _displayed_ correctly, but the underlying _data_ will be preserved (at least in Julia 0.7), and parsing HTML (which only looks at the ASCII characters, which are the same in Windows-1252 and UTF-8) should work. If you are just sending the non-ASCII data someplace else that expects Windows-1252, and you aren’t processing the non-ASCII text yourself, you can just pass it blindly through like this.
- Define a new `Windows1252String` type that represents Windows-1250 data directly; this is a lot of work and is probably not worth it.

Generally, I would recommend just converting the data to UTF-8, especially if the conversion cost is not performance-critical.
