# LaTeX table reader in Julia

**URL:** <https://discourse.julialang.org/t/latex-table-reader-in-julia/49067>\
**Category:** Data\
**Created:** [October 26, 2020, 7:47pm UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067 "2020-10-26T19:47:23Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sswatson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sswatson/32/5135_2.png) [@sswatson](https://discourse.julialang.org/u/sswatson)\
**Post date:** [October 26, 2020, 7:47pm UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/1 "2020-10-26T19:47:24Z")

</div>

Is there a LaTeX table _reader_ in Julia? I know there are several writers, but I haven’t been able to find something that takes the LaTeX source as input.

The reason I’m interested is that MathPix is really good at OCRing tables, but it outputs to LaTeX. (Example use case: I want to do some analysis with some data I find printed in a PDF.)

I know this would be really easy to throw together a basic version of, but I didn’t want to reinvent the wheel if it’s already out there.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 26, 2020, 8:05pm UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/2 "2020-10-26T20:05:07Z")

</div>

> [@sswatson](#):
>
> I know this would be really easy to throw together a basic version of

Ha. This was not _my_ first thought.

Not aware of anything. Is there an implementation in another language? You could start with that via some of the interop packages Julia has if so.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [October 26, 2020, 10:06pm UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/3 "2020-10-26T22:06:31Z")

</div>

Can you give an example of how a table would look?

---

<div class="post-metadata">

**Author:** ![sswatson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sswatson/32/5135_2.png) [@sswatson](https://discourse.julialang.org/u/sswatson)\
**Post date:** [October 26, 2020, 11:45pm UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/4 "2020-10-26T23:45:49Z")

</div>

To make sure the example is representative, I went to Google image search and OCRed one of the first tables I saw. Here was the output from MathPix:

```julia
\begin{aligned}
&\text { Table } 1.1 . \text { Nonlinear Model Results }\\
&\begin{array}{cccc}
\hline \hline \text { Case } & \text { Method#1 } & \text { Method#2 } & \text { Method#3 } \\
\hline 1 & 50 & 837 & 970 \\
2 & 47 & 877 & 230 \\
3 & 31 & 25 & 415 \\
4 & 35 & 144 & 2356 \\
5 & 45 & 300 & 556 \\
\hline
\end{array}
\end{aligned}

```

Here’s how it renders:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/3/c/3c1fb2ca572a59a855341bc03f0b174660eb8ced.png)

The part I’d be interested in is just the part between the `\begin{array}` and `\end{array}`.

@tbeason I’m starting to think you’re right about the difficulties here. At the most basic level, you’re just splitting into rows and then into entries. For the simplest examples, this would be trivial. But I’m thinking it’s likely to happen fairly often that MathPix outputs something that the naive algorithm isn’t prepared for. Handling that complexity in a graceful way is almost certainly more work than it’s worth.

My understanding is that there are LaTeX table readers in Python. That’s probably the right solution.

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [October 27, 2020, 12:20am UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/5 "2020-10-27T00:20:58Z")

</div>

Considering the different codes/environments one can use to generate a latex table, this is probably really difficult. I guess it would be marginally easy to develop something for mathpix only, as you have some gaurantee what the output code will include.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 27, 2020, 12:23am UTC](https://discourse.julialang.org/t/latex-table-reader-in-julia/49067/6 "2020-10-27T00:23:24Z")

</div>

You could try using Pandoc to convert to a simpler format (like markdown pipe tables) and then try to parse those (I think they’re basicslly CSVs).
