Turn complex documents into AI-ready data

Bring us the docs your pipeline struggles with. We’ll tune Hanji to the edge cases in your workflow.

COMPLIANCE
HIPAA + BAA
TRAINING
Never
SLA
Enterprise uptime SLAs

Built from production scale

70,000,000+

pages processed through the document pipeline that became Hanji.

Parse the pages. Extract the fields.

Grounded structure in, schema out, every value cited back to the source. Run them as one pipeline, or take Parse alone.

Parse

Any document in. Grounded structure out.

Referral packets, insurance claims, government forms, and scanned faxes come back as text, tables, and figures, with a bounding box on every element.

POST /v1/parse

Extract

Your schema in. Cited fields out.

Ungrounded fields come back null and flagged, so a guess never passes as fact.

Bounding boxes come straight from the parser, so every coordinate maps to real text.

POST /v1/extract/schema

Built for documents where a wrong field costs money.

Faxed referral packets, face sheets, and prior auths parse into clean tables and cited fields, so billers verify the source instead of re-keying scans. When fax quality causes misses, we tune on the exact failed documents. One customer improved from 71% to 92% text accuracy in under a week, at the same speed and cost.

Know where every field came from.

One pipeline reads, checks, and cites every field back to the exact spot it came from.

Grounded at read time

Every value leaves the page with its source region already attached, not aligned back after the fact. Vision models read text, tables, and layout in parallel.

Failures get caught, not passed through

Hanji catches blank, truncated, or garbled page reads and re-reads them with a heavier model. A bad read is flagged and withheld, never returned as clean data.

patientA. Rivera
payerAetna PPOmember_id8K2G-19Q
auth_statusYES
document → grounded read → structured fields

Fields cite their source

Every extracted value returns the exact quote, page, and bounding box it came from. Values that cannot be verified are flagged for review.

Parse accuracy benchmark

Highest score of 7 providers on two of four accuracy metrics. 400 human-verified pages. Full methodology

Hanji87.5%
Pulse82.7%
Reducto75.2%
AWS Textract74.1%
Docling72.5%
LlamaParse67.6%
Unstructured66.8%
Text accuracy
87.5%
+4.8 vs next best
Word-overlap F1
90.2%
+6.6 vs next best
Grounded accuracy
92.4%
-0.3 vs best
Layout IoU
78.1%
-7.9 vs best

400 human-verified pages · 7 document types · 95% confidence intervals. Learn more

Built for regulated documents

PHI under BAA, claims packets, public records. For teams whose documents drive eligibility, billing, and approvals.

HIPAA today, SOC 2 underway

PHI runs in production under a signed BAA today. The SOC 2 Type II audit is in progress.

Zero training

We never train on your data. Sync documents are processed in memory and never stored. Async uploads and results are deleted automatically after 3 days.

Throughput that holds at volume

Autoscaling sized for intake bursts, batch jobs up to a million pages, and uptime SLAs backed by a public status page.

Common questions

Something missing? Email hello@hanji.dev.

  • Sign up and you get 1,000 free credits. No card, no sales call, and full API access with keys straight from the dashboard. After the free tier it's pay as you go: parsing is 1 credit per page and schema extraction is 4 credits per page with the required parsing included, at a flat $0.003 per credit. If you'd rather see results on your own documents first, book a benchmark call.
  • A raw LLM returns values with nothing behind them and guesses when it's unsure. Hanji cites every value to its quote, page, and box, and returns null instead of guessing. An in-house pipeline hits the same wall: a confident wrong value looks exactly like a right one, with no flag to catch it.
  • On our benchmark of 400 human reviewed pages across seven document types, Hanji leads on text accuracy and word F1. Beyond that, accuracy is tuned per document type against a gold-labeled eval set, not a single global number. The honest answer for your documents is a measured one, so we'll run a free proof batch on your real files and show you field-level results before you commit. Book a proof batch.
  • Yes. HIPAA + BAA is available on request, and we have signed BAAs with healthcare customers in production. We never train on your data. The custom tier adds customer-managed encryption, configurable retention, and dedicated regions. Talk to us about your compliance requirements.
  • PDF, PPTX, DOCX, and images (PNG, JPEG, WebP, TIFF, HEIC/HEIF, BMP), with scanned PDFs and images handled automatically and OCR inline at no surcharge.
  • Yes. Send us 20-50 representative documents, or bring them to your benchmark call and we'll run the eval live. Results will be back within a few days. Healthcare and other regulated docs are handled under BAA on a private pipeline. Book a benchmark call.
  • Start on the free tier, then pay only for what you process: usage-based, with no seat fees. Higher volumes move to committed minimums and enterprise pricing scoped to your workload. See the pricing page for details, or book a call to scope your volume.
  • Yes, on the custom tier: dedicated regions, private networking, and self-hosting for strict security or data residency requirements, plus a negotiated credit rates, production SLAs, and a Slack channel with the engineering team. Contact us at hello@hanji.dev.
  • Synchronous requests support up to 500 pages or 150 MB. For larger batch jobs, our async endpoint supports up to 1M pages.

    Pages are parsed concurrently: on our public benchmark of 156 business PDFs and 13,671 pages, median latency was 10.7s per document, fastest among eight measured providers. The platform autoscales across requests, and we raise rate limits for high-volume customers. Use a benchmark call to size your target throughput.

Show us where it breaks

Bring us the docs your pipeline struggles with. We’ll work through the edge cases and tune Hanji for your workflow.