Document structure extraction without the heavyweight stack.
PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.
CPU only. No ML models. Runs in your browser, in Python, or as an API.
▶ Try it in your browser · Quick start · Benchmarks
30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (MP4)
Why papero
Getting the text out of a PDF is easy. Getting its structure back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
|
📖 Reading order Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside. |
▦ Real tables Ruled, borderless and LaTeX booktabs tables come back as rows and columns — multi-line cells included. Export to CSV or Excel. |
∑ Formulas Exponents, indices, stacked fractions and drawn root signs are read as mathematics — 5/12, 10√2, CO₂(g) — and written as LaTeX (\frac{5}{12}, \sqrt{2}), plus a cropped image of the formula.
|
|
📍 Position of everything Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it. |
🖼 Figures & charts Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together. |
📝 Back to Word Columns, alignment, indents, line spacing, bold runs and fonts are kept, so a .docx or HTML export looks like the original page — two columns stay two columns.
|
Also: accents drawn as separate glyphs in LaTeX PDFs (Computa¸ca˜o → Computação), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
Quick start
pip install papero-extract
from papero_extract import extract doc = extract("paper.pdf") print(doc.to_markdown())
Or skip the install: open the browser app, drop a PDF, export to the format you need.
More Python — tables, formulas, positions, images, options
from papero_extract import extract, extract_text doc = extract("paper.pdf", images=True) doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...] doc.formulas[0].latex # "E = mc^{2}" doc.figures[0].image.data # PNG bytes for block in doc.pages[0].blocks: # reading order, with positions print(block.type, block.bbox, block.text[:60]) doc.to_html() # keeps alignment and indents doc.to_dict() # the full JSON # Math inside paragraphs as LaTeX, in Markdown or plain text — found by what it is made of # (operators, functions, exponents, roots), even when it is set in the text font: doc.to_markdown(math="latex") # "Se $\operatorname{tg} x - \operatorname{cotg} x = 1$, então…" doc.to_text(math="latex") # "a) $1{,}035\cdot 10^{9}$ e $5{,}5\cdot 10^{7}$" extract("slides.pptx").to_markdown() # any format Apache Tika reads extract_text("contract.pdf").text # fastest: clean text only
| Option | Default | |
|---|---|---|
pages |
all | "1-3,5,10-" |
images |
False |
crop figures, tables and formulas to PNG |
tables / formulas |
True |
detection on/off |
ocr |
"auto" |
"auto" (scanned pages only), "force", "off" |
ocr_language |
"por+eng" |
Tesseract languages |
tika |
True |
False runs the layout engine alone (no Java) |
workers |
1 |
processes for long documents |
CLI
papero-extract extract paper.pdf -o paper.md --images # Markdown + images/ folder papero-extract extract paper.pdf -o paper.json # format from the extension papero-extract extract paper.pdf -f csv -o tables.csv # tables only papero-extract extract paper.pdf -p 1-5 -f html papero-extract extract paper.pdf --math latex -o p.md # math in the text as $…$ papero-extract extract paper.pdf --fast # clean text only papero-extract batch ./documents -o ./dataset # a whole folder, for RAG papero-extract serve --port 8000 # API + browser app
Batch & fidelity report — a folder of PDFs to a RAG dataset, and which ones to review
papero-extract batch ./documents -o ./dataset
dataset/
├── documents/ one .md and one .json per PDF (same sub-folders)
├── chunks.jsonl every chunk, cut at headings, tables kept whole
├── manifest.json per document: pages, tables, chunks, fidelity scores
└── fidelity/
├── report.json totals, signals and every document's issues
├── summary.html the same, to open in a browser
└── problematic/ one .json per document with warnings or errors
Each chunk knows where it came from, so a retrieval hit can be shown on the page:
{"id": "paper.pdf#12", "type": "table", "headings": ["4 Results"], "pages": [6, 6],
"blocks": ["p6-b3"], "text": "Table 2: Accuracy per model.\n\n| Model | Top-1 | …"}The report checks every document against its own PDF — no ground truth, so read it as where to look, not as accuracy:
| Signal | What is checked |
|---|---|
text |
the words PDFium reads on each page are all in the output |
reading_order |
no block is read after one below it in the same column |
tables |
every "Table N" caption has its table; each table is a clean grid |
figures |
every "Figure N" caption has its figure |
formulas |
each formula has LaTeX and no unmapped glyph |
The browser app runs the same checks on the PDF you drop: the fidelity figure sits next to the page count, and the Compare tab puts each page beside what was extracted from it, marking the PDF text that is not in the output and the blocks that failed a check.
from papero_extract import extract from papero_extract.batch import run_batch from papero_extract.chunks import chunk_document from papero_extract.fidelity import assess, reference_text run_batch("./documents", "./dataset", workers=8)["totals"] # {"documents": …, "ok": …, "warning": …, "error": …} doc = extract("paper.pdf") chunk_document(doc, document="paper.pdf", max_chars=1500) assess(doc, reference_text("paper.pdf")).issues # [Issue(code="table_not_detected", pages=[5], …)]
REST API & Docker
docker compose up # API + Apache Tika + Tesseract + browser app on :8000curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown" curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true" # blocks + positions
One endpoint, POST /v1/extract; interactive docs at /docs.
| Parameter | Default | |
|---|---|---|
mode |
structured |
structured (layout + Tika) or fast (text only) |
format |
json |
json, markdown, text, html, csv, zip |
pages |
all | 1-3,5,10- |
per_page |
false |
include pages, blocks and positions in the JSON |
images |
false |
crop figures, tables and formulas |
ocr |
auto |
auto, force, off |
Configuration through environment variables — see .env.example.
What comes out
Every block knows what it is and where it was:
{
"type": "table",
"bbox": [56.7, 294.8, 481.9, 374.2],
"rows": [["Model", "Accuracy"], ["Base", "0.81"]],
"caption": "Table 1: Comparison between models."
}| Output | Python · CLI · API | Browser app |
|---|---|---|
| Markdown, plain text, JSON | ✓ | ✓ |
| HTML (keeps alignment and indents) | ✓ | ✓ |
| CSV of the tables, ZIP with images | ✓ | ✓ |
Word .docx that keeps the page's look |
— | ✓ |
Excel .xlsx, one sheet per table |
— | ✓ |
Full JSON schema and block types
{
"schema": "pdf-text-api/document@1",
"engine": "tika+pdfium",
"page_count": 12,
"metadata": { "title": "...", "author": "...", "language": "en" },
"pages": [{
"number": 1, "width": 595.3, "height": 841.9,
"blocks": [{
"id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
"text": "Atestamos que a estudante ...",
"style": { "pt": 11.0, "font": "Arial", "bold": false },
"format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
"runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
}]
}]
}Block types: heading (with level), paragraph, list_item (with marker), table (with rows), figure, formula (with latex), caption, code, and — kept apart from the text — header, footer, page_number. Bounding boxes are [x0, y0, x1, y1] in points, origin at the top-left of the page.
Benchmarks
Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. papero · fast returns clean text; papero · structured also rebuilds reading order, tables, formulas and figures — 0 failures on 54 papers, 39 ms per page (median). Reproduce with benchmarks/.
| papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker | |
|---|---|---|---|---|---|---|
| License | MIT | AGPL | MIT | BSD | MIT | GPL |
| Needs ML models / PyTorch | no | no | no | no | yes | yes |
| Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
| Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
| Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
| Runs entirely in the browser | ✓ | — | — | — | — | — |
ML-based tools still win on very irregular layouts and complex math (matrices, aligned systems) — papero gives you the formula as approximate LaTeX and as an image so nothing is lost.
How it works
Two engines run on the same file at the same time:
- A layout engine on PDFium reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
- Apache Tika adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
Limitations
- Math: rebuilt from glyphs and strokes. Fractions, roots, exponents and indices are recognised; matrices, aligned systems and nested constructs come out linear (the cropped image is always there). In PDFs whose producer renumbered the glyphs of a math font, the Python engine can miss a symbol the browser engine reads by its glyph name.
- Word export keeps each page on its own page; where Word breaks lines differently, a dense page can run a few lines over onto an extra one.
- Borderless tables with very narrow gaps between columns can read as text.
- Scanned PDFs need OCR, which runs on the server path (Tesseract is in the Docker image).
- Word/Excel export is in the browser app for now.
Development
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd papero-pdf-text-extractor pip install -e ".[dev]" pytest -q # includes real-world regressions ruff check src tests && ruff format --check src tests npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out python -m http.server -d web # browser app at http://localhost:8000
src/papero_extract/ is the Python engine, API and CLI · web/ is the browser app (GitHub Pages) · tests/js/ checks the two engines agree · benchmarks/ downloads the dataset and draws the chart.
Contributing
Found a PDF papero gets wrong? That's the most useful issue you can open — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
If papero saves you time, a ⭐ helps other people find it.
Keywords: PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado
MIT © Beatriz Almeida · pip install papero-extract · import papero_extract