# PDF Parser and Reading API

**URL:** <https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797>\
**Category:** Data\
**Created:** [July 12, 2017, 8:34am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797 "2017-07-12T08:34:54Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [July 12, 2017, 8:34am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/1 "2017-07-12T08:34:54Z")

</div>

Hi All,

I am developing a PDF Library for doing some simple tasks like extracting texts and PDF file contents and attributes. The library is written in pure Julia (save some dependencies on some filter libraries). I am open to anyone interested in reviewing and contributing to it:

[https://github.com/sambitdash/PDFIO.jl](https://github.com/sambitdash/PDFIO.jl)

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![avik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/avik/32/17_2.png) [@avik](https://discourse.julialang.org/u/avik)\
**Post date:** [July 12, 2017, 11:00am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/2 "2017-07-12T11:00:13Z")

</div>

Hey Sambit,

Thanks for tackling this, a pure julia PDF parser will be very useful. I hope this can eventually lead to a writing library as well (but I’ll understand if that is not of particular interest to you). I’m also hoping that we can get some higher level tools on top, things like Tabula for example. I am very interested in this, and will play around with it, but unfortunately not sure how much time i’ll have to contribute seriously.

## Regards

Avik

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [July 12, 2017, 11:32am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/3 "2017-07-12T11:32:49Z")

</div>

Thanks Avik.

Tabula should not be as difficult, although every structured text processing in PDF is some form of heuristic as document structure is not mandated in PDFs. Acrobat had a table picker way back in 2002. So I am assuming it may not be hard to implement. Extraction of all the forms of text is definitely of my interest and will ensure APIs for the same are available. However, I may leave the subsequent heuristic development for table picking for someone to invest focused time and effort in that direction.

I will add an issue in the project for tracking this requirement.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [July 12, 2017, 11:45am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/4 "2017-07-12T11:45:12Z")

</div>

[https://github.com/sambitdash/PDFIO.jl/issues/2](https://github.com/sambitdash/PDFIO.jl/issues/2)

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [July 12, 2017, 3:01pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/5 "2017-07-12T15:01:12Z")

</div>

PDF parsing is a hideous problem.  
I wish you all the luck in the world.

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [July 12, 2017, 3:21pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/6 "2017-07-12T15:21:56Z")

</div>

@oxinabox

Very true!!! The biggest issue is PDF creators generate files that are non-compliant with the spec. Many a times you have to give the creator higher precedence over the spec based on your customer.

[http://www.stillhq.com/pdfdb/db.html](http://www.stillhq.com/pdfdb/db.html)

has a lot of such examples.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [July 19, 2017, 9:58am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/7 "2017-07-19T09:58:32Z")

</div>

Hi All,

I am now kind of finalizing the v1 of the APIs for the PDF library or the core of the PDF reader library. Here are the initial benefits of the library.

1. It will allow you to read through a PDF file and create objects which can be used for further access to the document.
2. It will also provide you the details of the content in every page and create a tree like data structure of PDF page contents which can be used know what is there in the PDF document.
3. The library has been tested with about 800+ text based files (12000+ pages) so fairly robust in text objects. And the parser is fairly robust and a bit non-tolerant as a standards based file is given higher emphasis.

However, next steps to extend the library requires specific domain where it will be used. For example, in the text extraction itself here are some standard challenges:

1. PDF text do not have reading order of character appearance. So text may appear as “aliuJ” with each character location in such a way printed such that the visual output is “Julia”.
2. Text and graphics directives can be interspersed. So you may get 5 different text objects as each character.
3. Since, fonts can be sub-setted “Julia” may be printed as (uvwxy) with gyph code of embedded font-51. One needs to query these judiciously with several logical smart reasoning to get the actual text.

Every such reasoning is subjective to the needs and interpretation of the developer/user and can be challenged with an alternate viewpoint. Hence, it’s important to keep the low level APIs simple and minimal such that any advanced development can be carried out on top of the minimal API set.

After some thoughts I realized I will rather keep the base APIs simple and minimal. Thus providing more flexibility to developers to develop more advanced solutions they need.

Of course there are a few areas in the basic APIs that are missing currently:

1. Enhancing the documentation of the library.
2. Support for encrypted PDF
3. Support for image filters. This has been knowingly avoided as most people may be using a third party API to render the final graphics. They could send the encrypted image in JPEG or JPX or LZW (TIFF, PNG, GIF) formats than decompressing and sending raw image to the rendering API.
4. Standardize the tree iterator with AbstractTrees APIs.
5. Develop what is needed as the adoption of APIs increase.

If you are all in agreement with my approach, I will register the PDFIO to Julia Package so that it’s available for general usage and testing.

Looking to hear from you soon.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [September 8, 2017, 8:12am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/8 "2017-09-08T08:12:50Z")

</div>

Update on the PDFIO API. It now has:

1. A full documentation of the APIs : [API Structure and Design · PDFIO](https://sambitdash.github.io/PDFIO.jl/docs/build/)
2. Has a text extraction API `pdPageExtractText(page)`
3. Complex page number support.
4. Supports unicode code extraction from font encoding as well as unicode CMap. (does not read into the font’s internal encoding)
5. Supports Adobe’s encoding for latin fonts.
6. Does not do any special handling for tagged PDFs but tagged PDFs may behave better as the creation order and reading order of document objects are similar.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [October 4, 2017, 6:57am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/9 "2017-10-04T06:57:47Z")

</div>

Update on the PDFIO API v.0.0.8:

Changes this release:

1. A new pdPageExtractText method is introduced which does a cleaner text conversion for complex PDFs including non-tagged PDFs.
2. Bug fixes  
Text conversions carried out on 25,000+ files.

The master untagged version has also some heuristics for text extraction when space character is simulated through text positioning. A few documents of 1000+ pages have been used for text extraction testing as well.

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [November 3, 2017, 2:38pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/10 "2017-11-03T14:38:29Z")

</div>

Update on the PDFIO API v.0.0.9:

Changes this release:

1. `pdPageExtractText` handles superscripts with enhanced heuristics.  
2, Space can be simulated from text positions.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [December 6, 2018, 8:50pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/11 "2018-12-06T20:50:09Z")

</div>

Is this working in Julia 1.0?  
Could I get the text out of a page of a PDF file?

---

<div class="post-metadata">

**Author:** ![jandehaan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jandehaan/32/6805_2.png) [@jandehaan](https://discourse.julialang.org/u/jandehaan)\
**Post date:** [December 6, 2018, 10:40pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/12 "2018-12-06T22:40:19Z")

</div>

> [@sambitdash](#):
>
> `pdPageExtractText`

Thank you for providing this PDF package.

I noticed that the function names that are part of the API don’t follow the Julia naming conventions, but other functions do. For example you chose a name like `pdPageExtractText` instead of something like `pd_page_extract_text`. Was this motivated by wanting to be consistent with PDF API naming conventions used elsewhere, or is this merely a historical accident?

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [December 7, 2018, 12:34am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/13 "2018-12-07T00:34:12Z")

</div>

The convention is very similar to what is used by Adobe’s PDF Library and many other libraries used in the industry in general.

Secondly, Julia does not have a convention for exported methods. Only exported methods in PDFIO follow this convention. Internal methods follow the underscore notation.

Regards,

Sambit

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [December 7, 2018, 12:35am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/14 "2018-12-07T00:35:05Z")

</div>

Yes. It works with 1.0.

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [December 7, 2018, 12:44am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/15 "2018-12-07T00:44:43Z")

</div>

I am trying to port an application from R and just need something like,

```julia
library(pdftools)
txt <- pdf_text(pdf = 'data.pdf') # Vector{String} for each page

```

How would that be with PDFIO.jl?

```julia
using PDFIO
doc = pdDocOpen("data.pdf")
page = pdDocGetPage(doc, 1) |>
    (doc -> IOBuffer() |>
        (io -> pdPageExtractText(io, page)))
read(page) # empty no data
pdDocClose(doc)

```

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [December 7, 2018, 1:47am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/16 "2018-12-07T01:47:11Z")

</div>

This link can give you some ideas.

[https://github.com/sambitdash/PDFIO.jl/issues/38](https://github.com/sambitdash/PDFIO.jl/issues/38)

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [December 7, 2018, 5:18am UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/17 "2018-12-07T05:18:21Z")

</div>

Perfect! Thanks.

Ended up using,

```julia
function getPDFText(pdf)
    doc = pdDocOpen(pdf)
    npage = pdDocGetPageCount(doc)
    for i ∈ 1:npage
        lines = pdDocGetPage(doc, i) |>
            (page -> pdPageExtractText(IOBuffer(), page)) |>
            (io -> readlines(IOBuffer(String(take!(io))))) |>
            magic
    end
end

```

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [April 1, 2019, 1:54pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/18 "2019-04-01T13:54:35Z")

</div>

I am getting `non-boolean (CosName) used in boolean context` with this [file](https://drive.google.com/file/d/1xsiPwiPsioRrEOB7gQTGG-xvK62YpBif/view?usp=sharing).

---

<div class="post-metadata">

**Author:** ![sambitdash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitdash/32/1377_2.png) [@sambitdash](https://discourse.julialang.org/u/sambitdash)\
**Post date:** [April 1, 2019, 2:26pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/19 "2019-04-01T14:26:42Z")

</div>

Give me a few days to investigate and get back. If you isolate the issue and feel free to submit a PR.

regards,

Sambit

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [April 1, 2019, 2:37pm UTC](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797/20 "2019-04-01T14:37:54Z")

</div>

I think it might require some more advanced PDF familiarity than what I have. R’s `pdftools::pdf_text` had no problem with the file. I was able to “fix it”, by running [ILovePDF compress](https://www.ilovepdf.com/compress_pdf) on the file and then it works just fine. The file was generated on a Macbook Pro using save as from a `cfm` file. Do reach out for more details or if you need the exact process to help track down the issue.

[Next page](https://discourse.julialang.org/t/pdf-parser-and-reading-api/4797.md?page=2)
