# How to extract data from pdf with two columns

**URL:** <https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008>\
**Category:** General Usage\
**Created:** [December 24, 2023, 4:05pm UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008 "2023-12-24T16:05:53Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Simone\_Gabbriellini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simone_gabbriellini/32/202218_2.png) [@Simone\_Gabbriellini](https://discourse.julialang.org/u/Simone_Gabbriellini)\
**Post date:** [December 24, 2023, 4:05pm UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/1 "2023-12-24T16:05:53Z")

</div>

I am trying to extract text from a pdf using PDFIO.jl.

Problem is that some slides have text on double columns, and the extraction with the provided example fails to reconstruct the text the proper way. For example using this:

```julia
doc = pdDocOpen(src)
docinfo = pdDocGetInfo(doc)
npage = pdDocGetPageCount(doc)
io = IOBuffer()
for i = 1:npage
    page = pdDocGetPage(doc, i)
    pdPageExtractText(io, page)
end
pdDocClose(doc)

```

on the slide below:

 ![Screenshot 2023-12-24 at 17.00.27](https://global.discourse-cdn.com/julialang/original/3X/9/7/97498f3522bfb241217c41e32fcbd694695c1554.png)

```julia
gives a poor performance:
● SAM: The solution ○ EU incubators / accelerators: ~1200 market size ○ EU VC: ~500 ● SOM: ○ IT incubators / ● TAM:○ 1.35M tech startups accelerators: ~250, worldwide ○ IT VCs: ~60 ○ ~7000 incubators / accelerators worldwide, ○ proxies to enter and counting 5000 startups ○ ~2500 VC ﬁrms, half of them on growth giano.rocks

```

so I am wondering if it would be possible to query the page for all the objects it contains and read them one by one, from left to right. But tbh I am quite lost… I know I can get the objects like:

```julia
page = pdDocGetPage(doc, 5)
elements = pdPageGetContentObjects(page)

```

any suggestion on how I can improve this to iterate over the objects and extract text? From the docs it looks like the only text extraction function wants a page and not an object, but maybe I missed something…

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 24, 2023, 4:17pm UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/2 "2023-12-24T16:17:24Z")

</div>

Would try using GPT4 or similar tools. I have had pretty good success giving it pdfs and then conducting tasks related to the files.

---

<div class="post-metadata">

**Author:** ![Simone\_Gabbriellini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simone_gabbriellini/32/202218_2.png) [@Simone\_Gabbriellini](https://discourse.julialang.org/u/Simone_Gabbriellini)\
**Post date:** [December 24, 2023, 6:30pm UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/3 "2023-12-24T18:30:01Z")

</div>

You mean asking gpt4 to write Julia code for the task of just ditch pdfio and use gpt4?

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 24, 2023, 7:03pm UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/4 "2023-12-24T19:03:40Z")

</div>

Just using GPT itself. I have given it unstructured text before and it works just fine to extract stuff. This was in ChatGPT, but I think the file upload API + GPT4 would probably work just as well. There is a thread floating around on Discourse here about Julia + AI packages, might be worth searching for in your case.

---

<div class="post-metadata">

**Author:** ![Simone\_Gabbriellini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simone_gabbriellini/32/202218_2.png) [@Simone\_Gabbriellini](https://discourse.julialang.org/u/Simone_Gabbriellini)\
**Post date:** [December 26, 2023, 9:49am UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/5 "2023-12-26T09:49:36Z")

</div>

thanks for your advice and for taking the time, but IMO ditching programming in favour of a paid service that does it for me kind of defies the very purpose of these forums.

might be me being old school… 🙂

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 26, 2023, 10:31am UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/6 "2023-12-26T10:31:39Z")

</div>

> [@Simone\_Gabbriellini](#):
>
> might be me being old school… 🙂

There seems to be a [similar open issue](https://github.com/sambitdash/PDFIO.jl/issues/17) in the package repo.

---

<div class="post-metadata">

**Author:** ![Simone\_Gabbriellini](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simone_gabbriellini/32/202218_2.png) [@Simone\_Gabbriellini](https://discourse.julialang.org/u/Simone_Gabbriellini)\
**Post date:** [December 26, 2023, 11:59am UTC](https://discourse.julialang.org/t/how-to-extract-data-from-pdf-with-two-columns/108008/7 "2023-12-26T11:59:09Z")

</div>

ahaha, you original message was a good burn…

thanks for the pointer, I missed that one
