# As simple as possible from the website, extract the raw text?

**URL:** <https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697>\
**Category:** General Usage\
**Created:** [January 30, 2018, 3:52pm UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697 "2018-01-30T15:52:24Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [January 30, 2018, 3:52pm UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697/1 "2018-01-30T15:52:24Z")

</div>

As simple as possible from the website, extract the raw text ?  
Paul

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [January 30, 2018, 11:44pm UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697/2 "2018-01-30T23:44:11Z")

</div>

Simplified from [https://github.com/oxinabox/DataDepsGenerators.jl/blob/master/src/utils.jl](https://github.com/oxinabox/DataDepsGenerators.jl/blob/master/src/utils.jl)

```julia
using Gumbo

getpage(url) = parsehtml(String(read(download(url))))
text_only(doc::HTMLDocument) = text_only(doc.root)
text_only(frag) = join([text(leaf) for leaf in Leaves(frag) if leaf isa HTMLText], " ")

get_page_text(url) = text_only(getpage(url))

```

---

<div class="post-metadata">

**Author:** ![sdmcallister](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdmcallister/32/2324_2.png) [@sdmcallister](https://discourse.julialang.org/u/sdmcallister)\
**Post date:** [January 31, 2018, 12:14am UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697/3 "2018-01-31T00:14:40Z")

</div>

Nice–Thanks.

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [January 31, 2018, 7:48am UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697/4 "2018-01-31T07:48:51Z")

</div>

> text\_only(doc.root)  
> ERROR: UndefVarError: text\_only not defined

Version 0.6.0 (2017-06-19 13:05 UTC)  
Official [http://julialang.org/](http://julialang.org/) release  
x86\_64-w64-mingw32

using Gumbo  
using DataDeps  
using DataDepsGenerators

Paul

W dniu 2018-01-31 o 00:49, Lyndon White pisze:

---

<div class="post-metadata">

**Author:** ![programista](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/programista/32/5372_2.png) [@programista](https://discourse.julialang.org/u/programista)\
**Post date:** [February 9, 2018, 4:16pm UTC](https://discourse.julialang.org/t/as-simple-as-possible-from-the-website-extract-the-raw-text/8697/5 "2018-02-09T16:16:30Z")

</div>

ERROR: MethodError: no method matching tag(::Gumbo.HTMLText)  
Closest candidates are:  
tag(::Gumbo.HTMLElement{T}) where T at  
C:\Users\PC.julia\v0.6\Gumbo\src\manipulation.jl:6

for:  
url=“[http://www.rp.pl](http://www.rp.pl)” (and ohers )  
doc=parsehtml(String(read(download(url))))

Some new idea?  
Paul

W dniu 2018-01-31 o 01:20, Sdmcallister pisze:
