# PDFIO pdPageExtractText

**URL:** <https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781>\
**Category:** New to Julia\
**Created:** [March 10, 2020, 12:50am UTC](https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781 "2020-03-10T00:50:06Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![BMval](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bmval/32/7647_2.png) [@BMval](https://discourse.julialang.org/u/BMval)\
**Post date:** [March 10, 2020, 12:50am UTC](https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781/1 "2020-03-10T00:50:06Z")

</div>

Does anybody work with PDFIO.jl?  
I need to extract text from pdf, due to limiting of my knowledge, I don’t understand how to use the method to save text to string, which is printing to REPL.

```julia
 doc = pdDocOpen("test.pdf")
 page = pdDocGetPage(doc, 1)
 dPageExtractText(stdout, page)

```

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [March 10, 2020, 1:20am UTC](https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781/2 "2020-03-10T01:20:42Z")

</div>

the first argument to `dPageExtractText` is `stdout` indicating that the output is printed. If you want to capture the output, you may try something like this

```julia
io = IOBuffer()
dPageExtractText(io, page)
String(take!(io))

```

note, that I have not tried this, but something like that should work.

---

<div class="post-metadata">

**Author:** ![BMval](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bmval/32/7647_2.png) [@BMval](https://discourse.julialang.org/u/BMval)\
**Post date:** [March 10, 2020, 3:09am UTC](https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781/3 "2020-03-10T03:09:07Z")

</div>

Thank you .  
only, it should be String() not string()

---

<div class="post-metadata">

**Author:** ![cobrp](https://avatars.discourse-cdn.com/v4/letter/c/977dab/32.png) [@cobrp](https://discourse.julialang.org/u/cobrp)\
**Post date:** [February 4, 2025, 12:30pm UTC](https://discourse.julialang.org/t/pdfio-pdpageextracttext/35781/4 "2025-02-04T12:30:04Z")

</div>

I use this function:

function getPDFText(src, out)  
# handle that can be used for subsequence operations on the document.  
doc = pdDocOpen(src)

```
# Metadata extracted from the PDF document. 
# This value is retained and returned as the return from the function. 
docinfo = pdDocGetInfo(doc)
open(out, "w") do io
    # Returns number of pages in the document       
    npage = pdDocGetPageCount(doc)
    for i=1 : npage
        # handle to the specific page given the number index. 
        page = pdDocGetPage(doc, i)
        # Extract text from the page and write it to the output file.
        pdPageExtractText(io, page)
    end
end

```

I found it on the internet. My problem is that it take always the first line of the next page. The PDFtk does a better job.
