TECHNOLOGY

New Open-Source Tool Papero Extracts Structured Text From PDFs Without Machine Learning

Papero is a free, MIT-licensed PDF extraction tool that recovers reading order, tables, formulas and layout using geometry alone, running entirely on a CPU in the browser, Python or via API.

Laptop screen showing a PDF document with colored bounding boxes outlining text, tables and formula blocks in a browser appTECHNOLOGY

Image: IntraGoals Media · Uploaded by IntraGoals — usage rights confirmed

A new open-source project called papero is offering developers a lightweight alternative to machine-learning-heavy tools for extracting structured content from PDF files. Released under the MIT license and shared on Hacker News by developer Beatriz Almeida Fernandes, the tool converts PDFs into Markdown, JSON, Word or Excel formats while preserving reading order, tables, formulas, figures and the precise position of every block of content on the page.

What sets papero apart from many competing tools is what it doesn't use: no machine learning models, no PyTorch, no GPU. Instead, it relies on plain geometry to reconstruct a document's structure, reading glyph positions, fonts, rules and images directly from the PDF to rebuild columns, tables, formulas and lists through what the project describes as a column-aware XY-cut algorithm. That approach, the developer says, keeps the tool fast even on a laptop CPU.

The project addresses a common frustration in document processing. Pulling raw text out of a PDF is trivial, but recovering its structure — which column should be read first, which lines form a table, where a mathematical formula sits — is what determines whether the output is actually useful for applications like retrieval-augmented generation (RAG), search indexing or feeding documents to large language models.

Papero runs in three ways: as a browser app where users can drag in a PDF and inspect it without the file ever leaving their machine, as a Python package installable via pip, or as a REST API with Docker support. The browser version runs the same extraction algorithm ported to JavaScript using pdf.js, while the Python and API versions use a layout engine built on PDFium paired with Apache Tika, which adds metadata handling, OCR through Tesseract for scanned pages, and support for additional formats including DOCX, PPTX, XLSX, EPUB and HTML.

The tool handles some specific edge cases that trip up other extractors, according to the project page. It reconstructs formulas such as fractions, exponents and roots as LaTeX along with a cropped image of the formula, reassembles accented characters that some LaTeX-generated PDFs split into separate glyphs, and discards invisible white text sometimes inserted by form-generation software.

According to benchmarks published by the developer, papero's structured mode processed dense, multi-column arXiv papers with tables, formulas and figures at a median of 39 milliseconds per page with zero failures across 54 test papers, run entirely on a laptop CPU with no GPU. The developer's comparison table shows papero's average extraction time landing at 543 milliseconds per PDF in structured mode, slower than PyMuPDF's 97 milliseconds but dramatically faster than ML-based tool Docling, which the benchmark clocked at 82.9 seconds per document.

The developer acknowledges papero has limits. Machine-learning-based tools reportedly still outperform it on highly irregular layouts and complex mathematical notation, such as matrices or aligned equation systems. In those cases, papero still outputs an approximate LaTeX rendering alongside a cropped image of the formula so the content isn't lost entirely.

The project is open for contributions, with the developer specifically inviting users to submit PDFs that papero processes incorrectly, along with notes on what the expected output should have been.

Sources and further readingGitHub - beatrizalmeidaf/papero-pdf-text-extractor: Fast, open-source PDF text extraction API. Files never stored. ↗
ABOUT THE DESK

IntraGoals News Desk

IntraGoals reports on important changes in technology and work. We check each story for clear writing, trusted sources and useful information before it is published.

KEEP READING

Latest from IntraGoals.

All latest stories ↗
New Open-Source Tool Papero Extracts Structured Text From PDFs Without Machine Learning | IntraGoals