Full 156-document corpus breakdown
Supporting data for extract-bench across 156 business PDFs and 13,671 pages. This page measures speed, coverage, and provider-agreement consensus; consensus F1 is not human-labeled accuracy, so the 400-page gold set remains the primary accuracy claim.
p50 wall-clock
151/156 documents in selected tags
Upper-left is better: faster latency with higher consensus against the majority-provider reference.
| Provider | Mode | p50 latency | Consensus F1 | Pages per second | Success |
|---|---|---|---|---|---|
| AWS Textract | LAYOUT + TABLES | 29.5s | 99.7% | 1 | 155/156 |
| Azure DI | prebuilt-layout | 18.2s | 99.4% | 1.2 | 154/156 |
| Hanji | hosted | 10.7s | 98.8% | 2 | 151/156 |
| Google Doc AI | batch layout | 75.4s | 99.6% | 0.5 | 155/156 |
| LlamaParse | agentic | 29.1s | 97.7% | 0.8 | 154/156 |
| Reducto | standard parse | 23.8s | 99.0% | 1.2 | 155/156 |
Exact measurements for the selected documents.
| Provider | Mode | Success | Consensus F1 | p50 latency | Pages/sec | BBox |
|---|---|---|---|---|---|---|
| hosted | 151/156 | 98.8% | 10.7s | 2 | 100% | |
| Azure DI | prebuilt-layout | 154/156 | 99.4% | 18.2s | 1.2 | 100% |
| standard parse | 155/156 | 99% | 23.8s | 1.2 | 100% | |
| agentic | 154/156 | 97.7% | 29.1s | 0.8 | 100% | |
| LAYOUT + TABLES | 155/156 | 99.7% | 29.5s | 1 | 100% | |
| Google Doc AI | batch layout | 155/156 | 99.6% | 75.4s | 0.5 | 100% |
Rasterized docs ship as page-image PDFs (no text layer). They test the OCR / layout path, not text-layer extraction. Treat them as a sibling-class to the scanned-doc set, not equivalent to born-digital business PDFs.
- 156 documents across representative classes.
- Default rankings require at least 85% document success and non-zero source bbox coverage.
- The corpus spans tiny born-digital, academic, financial filings, compliance / regulatory, legal, healthcare, insurance, government forms, technical specs, manuals, slides, image-heavy / magazines, multilingual, regression cases, and DocLayNet rasterized layouts.
- DocLayNet docs are stitched from rasterized page-images (no text layer). They test the OCR / layout path, not text-layer extraction. The page tags them is_rasterized.
- Consensus F1 is computed against tokens emitted by a majority of successful providers for each document.