# Effective Text Extraction from Documents (PDFs)

**URL:** <https://discourse.julialang.org/t/effective-text-extraction-from-documents-pdfs/54889>\
**Category:** General Usage\
**Tags:** question, strings, data, nlp, etl\
**Created:** [February 9, 2021, 2:24am UTC](https://discourse.julialang.org/t/effective-text-extraction-from-documents-pdfs/54889 "2021-02-09T02:24:35Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![TheCedarPrince](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thecedarprince/32/17323_2.png) [@TheCedarPrince](https://discourse.julialang.org/u/TheCedarPrince)\
**Post date:** [February 9, 2021, 2:24am UTC](https://discourse.julialang.org/t/effective-text-extraction-from-documents-pdfs/54889/1 "2021-02-09T02:24:35Z")

</div>

Hi all,

Question for those involved with text extraction pipelines (or ETL pipelines for that matter):

In Julia, what works best for you when doing text extraction from PDFs?

In my case, I am simplifying the problem to look at PDFs that are single column in form and are written in English with either none or minimal images. Currently trying out both [Taro.jl](https://github.com/aviks/Taro.jl) and [PDFIO.jl](https://github.com/sambitdash/PDFIO.jl) - so far I have found Taro.jl a bit easier to work with to extract content. However, both packages struggle with white space in the form of actual spaces between sentences or new line characters so it seems where I end up with long delineated concatenated strings of words…

Any tips/tricks on text extraction that might help better with my processing?

Thank you!

~ tcp 🌳

---

<div class="post-metadata">

**Author:** ![contradict](https://avatars.discourse-cdn.com/v4/letter/c/ac91a4/32.png) [@contradict](https://discourse.julialang.org/u/contradict)\
**Post date:** [February 9, 2021, 7:28pm UTC](https://discourse.julialang.org/t/effective-text-extraction-from-documents-pdfs/54889/2 "2021-02-09T19:28:29Z")

</div>

Not Julia, but a friend of mine recently did this and found that using Inkscape to convert to SVG and then parsing the SVG worked well. I think this is where they started: [https://github.com/scraperwiki/pdf2svg/blob/master/pdf2svg.sh](https://github.com/scraperwiki/pdf2svg/blob/master/pdf2svg.sh)

---

<div class="post-metadata">

**Author:** ![mbaz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbaz/32/17295_2.png) [@mbaz](https://discourse.julialang.org/u/mbaz)\
**Post date:** [February 9, 2021, 10:48pm UTC](https://discourse.julialang.org/t/effective-text-extraction-from-documents-pdfs/54889/3 "2021-02-09T22:48:29Z")

</div>

There’s also [pdftotext](https://www.xpdfreader.com/pdftotext-man.html).
