Papero extracts PDF structure with plain geometry, no ML models
Lightweight PDF parser with layout, tables, formulas and bounding boxes

Papero is an open-source PDF parser that recovers reading order, tables, formulas, figures, and bounding boxes using only CPU and geometric algorithms—no ML models or PyTorch. It runs in the browser, as a Python package, or as a REST API, and processes dense arXiv papers at 39 ms per page with zero failures on 54 papers. Output includes Markdown, JSON, Word, and Excel with full layout fidelity.
Getting the text out of a PDF is easy. Getting its structure back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs.