# Way to transform HTML into Text?

**URL:** <https://discourse.julialang.org/t/way-to-transform-html-into-text/35434>\
**Category:** General Usage\
**Created:** [March 3, 2020, 12:01am UTC](https://discourse.julialang.org/t/way-to-transform-html-into-text/35434 "2020-03-03T00:01:22Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![djsegal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djsegal/32/13752_2.png) [@djsegal](https://discourse.julialang.org/u/djsegal)\
**Post date:** [March 3, 2020, 12:01am UTC](https://discourse.julialang.org/t/way-to-transform-html-into-text/35434/1 "2020-03-03T00:01:22Z")

</div>

It seems like [Gumbo.jl](https://github.com/JuliaWeb/Gumbo.jl) does a good job of parsing HTML, but there doesn’t seem to be an easy way to extract the text from it (akin to beautiful soup in python).

Has anyone encountered this problem? How did you turn a whole `HTMLDocument` into a text string?

---

<div class="post-metadata">

**Author:** ![djsegal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djsegal/32/13752_2.png) [@djsegal](https://discourse.julialang.org/u/djsegal)\
**Post date:** [March 3, 2020, 12:43am UTC](https://discourse.julialang.org/t/way-to-transform-html-into-text/35434/2 "2020-03-03T00:43:24Z")

</div>

Cobbled together this code that kind of does what I want?

```julia
using Gumbo
using AbstractTrees

import Gumbo.text

function text(cur_doc::HTMLDocument)
    string_parts = []

    for elem in PreOrderDFS(aaa.root) 
        isa(elem, HTMLText) || continue
        push!(string_parts, Gumbo.text(elem))
    end

    return join(string_parts, " ")
end

```
