Bring us the docs your pipeline struggles with. We’ll tune Hanji to the edge cases in your workflow.
Built from production scale
pages processed through the document pipeline that became Hanji.
Grounded structure in, schema out, every value cited back to the source. Run them as one pipeline, or take Parse alone.
Parse
Referral packets, insurance claims, government forms, and scanned faxes come back as text, tables, and figures, with a bounding box on every element.
Extract
Ungrounded fields come back null and flagged, so a guess never passes as fact.
Bounding boxes come straight from the parser, so every coordinate maps to real text.
Faxed referral packets, face sheets, and prior auths parse into clean tables and cited fields, so billers verify the source instead of re-keying scans. When fax quality causes misses, we tune on the exact failed documents. One customer improved from 71% to 92% text accuracy in under a week, at the same speed and cost.
One pipeline reads, checks, and cites every field back to the exact spot it came from.
Every value leaves the page with its source region already attached, not aligned back after the fact. Vision models read text, tables, and layout in parallel.
Hanji catches blank, truncated, or garbled page reads and re-reads them with a heavier model. A bad read is flagged and withheld, never returned as clean data.
Every extracted value returns the exact quote, page, and bounding box it came from. Values that cannot be verified are flagged for review.
Highest score of 7 providers on two of four accuracy metrics. 400 human-verified pages. Full methodology
400 human-verified pages · 7 document types · 95% confidence intervals. Learn more
PHI under BAA, claims packets, public records. For teams whose documents drive eligibility, billing, and approvals.
PHI runs in production under a signed BAA today. The SOC 2 Type II audit is in progress.
We never train on your data. Sync documents are processed in memory and never stored. Async uploads and results are deleted automatically after 3 days.
Autoscaling sized for intake bursts, batch jobs up to a million pages, and uptime SLAs backed by a public status page.
Something missing? Email hello@hanji.dev.
Synchronous requests support up to 500 pages or 150 MB. For larger batch jobs, our async endpoint supports up to 1M pages.
Pages are parsed concurrently: on our public benchmark of 156 business PDFs and 13,671 pages, median latency was 10.7s per document, fastest among eight measured providers. The platform autoscales across requests, and we raise rate limits for high-volume customers. Use a benchmark call to size your target throughput.
Bring us the docs your pipeline struggles with. We’ll work through the edge cases and tune Hanji for your workflow.