400 pages from public datasets and documents, human-verified region by region and run through 7 providers across seven document types.
Hanji leads two of four columns. Means with 95% confidence intervals, sorted by text accuracy.
| Provider | Text accuracy | Word-overlap F1 | Grounded accuracy | Layout IoU |
|---|---|---|---|---|
87.5% 85.6%–89.4% | 90.2% 88.8%–91.5% | 92.4% 91.0%–93.6% | 78.1% 76.1%–80.2% | |
82.7% 80.9%–84.7% | 83.6% 82.0%–85.1% | 92.7% 91.6%–93.8% | 80.1% 78.3%–81.8% | |
75.2% 73.1%–77.5% | 77.1% 74.8%–79.2% | 86.0% 84.6%–87.3% | 86.0% 84.4%–87.5% | |
74.1% 71.1%–77.0% | 76.3% 73.5%–79.1% | 79.0% 76.3%–81.4% | 46.9% 44.9%–49.0% | |
72.5% 69.7%–75.2% | 77.7% 75.1%–80.5% | 82.6% 80.2%–84.9% | 74.4% 72.0%–76.7% | |
67.6% 64.7%–70.3% | 69.7% 67.4%–72.1% | 65.9% 63.1%–68.6% | 48.6% 46.0%–51.1% | |
66.8% 64.0%–69.6% | 63.9% 60.7%–67.0% | 70.4% 67.3%–73.3% | 83.0% 81.0%–85.0% |
Grounded accuracy by document type. Hanji holds the highest floor in the field: even its weakest type scores 86.6%, while every other provider drops below that on at least one type. Specialists edge ahead on individual cuts; none stays this consistent across all seven.
Hanji leads this cut by 1.3 points over Pulse.







Four metrics, each with its scoring rule. Every score is a document‑level mean with a bootstrap confidence interval.
For each labeled region we gather the predicted boxes overlapping its location (page-sized dumps are excluded), stitch their text in reading order, and measure how much of the region's text is recoverable in that local context — when the text is contained, otherwise the best-matching window scored by . The metric averages this recall over all regions, with a region that got no overlapping prediction scoring 0; extra neighboring text from a coarse box is not penalized, so word-, line-, and paragraph-box outputs are scored fairly.
Character fidelity over the whole page in each provider's emitted reading order: , where is the Levenshtein distance between prediction and ground truth over the longer length. Markdown, case, and Textract LAYOUT table scaffolding are normalized first.
Bag-of-words overlap, independent of order: the harmonic mean of word precision and recall , . It surfaces which words were captured versus dropped or hallucinated.
How well the predicted boxes cover the real content blocks: intersection-over-union of the predicted () and verified () box masks rasterized on a 1000×1000 grid, , independent of the text inside.
All 400 pages span seven document types: forms & invoices, tables & financial, multi-column, multilingual, image-heavy, presentations, and scanned & degraded. They are drawn from public OCR datasets (CORD, SROIE), arXiv pages, lecture slides, SEC 10-K filings, and synthetic multilingual renders. Every page was human-verified region by region against its source. Each score is a document-level mean with a 95% confidence interval from bootstrap resampling over pages. Test pages are held out from training.
The larger consensus benchmark asks a different question: across 156 business PDFs and 13,671 pages, who stays fast while remaining competitive with the provider-agreement reference? Corpus mix: financial reports, compliance audits, legal contracts, healthcare and insurance packets, government forms, technical specs, manuals, slides, image-heavy pages, multilingual docs, and scanned/OCR cases.
| Provider | Mode | p50 latency | Consensus F1 | Pages per second | Success |
|---|---|---|---|---|---|
| AWS Textract | LAYOUT + TABLES | 29.5s | 99.7% | 1 | 155/156 |
| Azure DI | prebuilt-layout | 18.2s | 99.4% | 1.2 | 154/156 |
| Hanji | hosted | 10.7s | 98.8% | 2 | 151/156 |
| Google Doc AI | batch layout | 75.4s | 99.6% | 0.5 | 155/156 |
| LlamaParse | agentic | 29.1s | 97.7% | 0.8 | 154/156 |
| Reducto | standard parse | 23.8s | 99.0% | 1.2 | 155/156 |
Bring us the docs your pipeline struggles with. We’ll work through the edge cases and tune Hanji for your workflow.