Naktor
BENCHMARK · JUL 2026

10 document-extraction systems, actually measured

Quality via the official OmniDocBench harness and a dual LLM judge, real cost (not list price) and latency, over the same 100 pages. Four results the industry's marketing won't tell you.

By Hektor Jacynycz, CTO and co-founder at Naktor23 July 2026

/1,000 pages — podium quality at the best price
$0.99
table TEDS — the best quality on the board
90.9
cheaper than cloud OCR, at better quality
on-prem — your documents never leave your infra
100%

Naktor · 100 OmniDocBench pages · reproducible

EXECUTIVE SUMMARY

Four things the numbers say and marketing doesn't

We measured quality, cost and latency over the same public sample. This is what survived adversarial verification.

01

pricier, worse tables

Mistral OCR 4's premium doesn't buy quality

At $4/1,000 (4× naktor-extractor) it ties on text, but loses 15 TEDS points on tables (75.8 vs 90.9) and doubles the formula distance. Going from OCR 3 to OCR 4 buys no measurable gain on this corpus.

02

0 / 3

core metrics won vs naktor

Unlimited-OCR's hype doesn't survive the harness

Marketed by Baidu as state of the art, it loses to or ties with naktor-extractor on text, tables and formulas. It only wins clearly on handwriting (0.22 vs 0.34): real merit, far from «best at everything».

03

0.0 → 84.7

TEDS, same model

Serving mode matters more than the model

The same PaddleOCR-VL goes from unusable (TEDS 0.0, served «bare») to second on tables (84.7, official pipeline). Any comparison that doesn't document serving is suspect.

04

on-prem

owns the frontier

The best quality-per-euro never leaves your infra

naktor-extractor delivers top-table quality at $0.99/1,000, on your own infrastructure. No cloud API comes close to that ratio —and every one requires shipping your documents out.

COST × QUALITY

More quality. Less cost.

Each point is a system over the 100 pages. The higher up, the more quality; the further left, the lower the cost. The top-left corner is the good one — and naktor-extractor is almost alone in it.

More quality. Less cost. — Text quality (100 − 100·edit distance)7883889398$0.5$1$2$4$8$/1,000 pages (log scale)Text quality (100 − 100·edit distance)naktor-extractor$0.99 · podium qualityPaddleOCR-VLUnlimited-OCRMistral OCR 4Mistral OCR 3DeepSeek-OCRolmOCR-2Docling
naktor-extractoron-prem — documents never leave your infracloud API

For cloud APIs and naktor-extractor the axis is the price; for the open-source systems it is the cost of self-hosting them (which requires your own infrastructure and operations). On-prem includes the CPU tool (Docling).

RESULTS

Every system, column by column

Sort by any column. The naktor-extractor row is highlighted; ↓ = lower is better, ↑ = higher is better.

systemgovernance
naktor-extractoron-prem0.0890.03690.90.10089.981.0$0.995.6
PaddleOCR-VL (official pipeline)on-prem0.0660.04684.70.09085.181.7$4.9910.9
Unlimited-OCR (Baidu)on-prem0.1020.06790.00.12087.082.2$0.747.7
Mistral OCR 4cloud0.0900.073*75.80.20889.183.6$4.002.2
Mistral OCR 3cloud0.0960.033*77.90.20385.178.1$2.001.4
DeepSeek-OCRon-prem0.1410.06673.50.27279.777.0$0.674.0
olmOCR-2 7Bon-prem0.1490.09167.20.21781.479.5$2.1532.1
DoclingCPU0.1890.10663.20.93675.470.2$1.453.1
(ablation) PaddleOCR-VL bareon-prem0.1970.1840.00.7729.6

* See Finding 3: a single page distorts Mistral's English text score.

TEDS: table-structure similarity, 0–100 (higher is better). p50: per-page median. Judges A/B: two multimodal LLMs from different families.

The cost-column bars use a log scale.

DECISION GUIDE

Which engine to pick for your case

There is no «best OCR»: there is a best one for your dominant KPI and your constraints. This guide follows directly from the measured numbers.

Dominant KPI · $/page with podium quality

Bulk archive digitization, corporate RAG, back-office

Pick

naktor-extractor

$0.99/1,000, TEDS 90.9 (best), text 0.089; on your own infrastructure.

Dominant KPI · Data never leaves your infra

Sensitive data (banking, health, legal, government)

Pick

naktor-extractor (deploys on your infrastructure)

Every cloud API is disqualified on governance before price is even discussed.

Dominant KPI · Cost on mixed corpora

Real-world mix of digital + scanned PDFs

Pick

naktor-extractor

Digital documents cost far less than scanned ones: you only pay a premium for what's genuinely hard.

Dominant KPI · Minimum edit distance

Maximum text/formula fidelity, cost secondary

Pick

PaddleOCR-VL (official pipeline)

Best text (0.066) and formulas (0.090) — at $4.99/1,000 (5× naktor's price) and single-threaded throughput.

Dominant KPI · Handwriting quality

High volume of handwriting

Pick

Unlimited-OCR

0.22 vs naktor-extractor's 0.34 on handwriting; cheap to self-host ($0.74/1,000).

Dominant KPI · Per-page latency

Interactive few-page flow, user waiting

Pick

Mistral OCR 4 (or OCR 3 at half price)

2.2s/page via API with no infra to manage; top fidelity judge (83.6). OCR 3 is near-equal for $2.

Dominant KPI · CPU simplicity

No GPU, minimal budget, no formulas

Pick

Docling

$1.45/1,000 on CPU, acceptable text (0.189); unusable for formulas (0.936).

Post-benchmark note: with the fix from Finding 6, the «maximum fidelity» row narrows — naktor-extractor now measures 0.064 on text (better than PaddleOCR-VL's 0.066); PaddleOCR-VL's edge is confined to formulas.

THE FULL PAPER

Method, findings and limits — uncut

The summary above comes from here. This is the part you can quote, audit and rebut.

Why this benchmark exists

Every few weeks a launch («the definitive OCR») claims state of the art with self-evaluated numbers and no real cost. naktor-extractor —our extraction service— has real money riding on the truth. This benchmark answers with method what the hype answers with adjectives: how much quality does each dollar buy, and under what data-governance conditions?

Methodology

Corpus. 100 pages sampled with stratification (fixed seed, script published) from OmniDocBench (CVPR 2025), the standard benchmark Mistral, Baidu and DeepSeek report against —revision pinned (aa1ee96, v1.6 content, 1,651 pages). Strata: 12 handwritten, 10 degraded scans, 20 dense tables, 15 formula-heavy, 13 hard layouts, 15 figure-rich, 15 dense text; 57 pages contain English, 41 Chinese. The sample deliberately over-weights the *hard* subsets: absolute values run lower than the public leaderboard — what matters is the comparison under identical conditions.

Systems and serving. Cloud: Mistral OCR 4 and 3 (API). Self-hosted on cloud GPUs with each system's official recipe: Unlimited-OCR 3B, DeepSeek-OCR, PaddleOCR-VL 0.9B (official pipeline; the bare mode is reported separately as an ablation), olmOCR-2 7B. CPU: Docling and Marker on a local M-series machine (sequential). naktor-extractor ran in its production configuration.

Quality. Two independent layers. (1) The official OmniDocBench harness (quick_match, GT filtered to the 100 pages): text edit distance, table TEDS, formula distance and reading order. (2) A dual multimodal LLM judge that sees the image and the markdown and scores completeness, fidelity, structure and hallucination, with two judges from different families so no verdict rests on a single model. Layer (1) is strict about structure; layer (2) measures whether the *information* survives even when formatting differs. Divergences get investigated, not averaged.

Cost. Measured, not list: APIs via actual per-page billing; self-hosted via wall-clock of a warm concurrent 100-page batch × the cloud GPU hourly rate, at the maximum concurrency each serving supports. Cold start is measured and reported separately. Local CPU: time × an equivalent cloud container rate. The figure for naktor-extractor is its price, not its cost.

Adversarial verification. Every anomaly was attacked before publication, with the pages and diffs in hand. Raw artifacts exist locally for audit; they are not republished because the corpus copyright forbids it (see Limitations).

Verified findings

1. Mistral OCR 4's price premium doesn't buy quality. At $4/1,000 (2× OCR 3, 4× naktor-extractor), OCR 4 ties on text (0.090 vs 0.089), loses 15 TEDS points on tables and doubles the formula distance. Its real strengths are per-page API latency (2.2s p50) and the highest fidelity score from the second judge (83.6) — a fine product for interactive few-page flows; for bulk digitization its price can't be defended against the on-prem frontier.

2. The Unlimited-OCR case: the hype doesn't survive the harness. Baidu's launch claims state of the art on OmniDocBench. On our hard sample, using its official vLLM recipe, it loses to or ties with naktor-extractor on text, tables and formulas, winning clearly only on handwriting (0.22 vs 0.34) — real merit, far from «best at everything». Its architecture (3B MoE, constant KV) does make it cheap to self-host ($0.74/1,000).

3. Metrics lie too, unless audited: one artifact we almost published. OCR 3 looked better than OCR 4 on English (0.033 vs 0.073); 93% of the gap was ONE page —a table of contents OCR 4 rendered as a table while the GT annotates it as plain text; edit distance punishes the markup, not the reading. Without that page: a tie (0.034 vs 0.037). We publish it because a benchmark that hides its own false positives deserves no trust.

4. Serving is part of the model. PaddleOCR-VL goes from 0.197/TEDS 0.0 (bare VLM, generic prompt) to 0.066/TEDS 84.7 (official pipeline with layout + per-element prompts): the same checkpoint, 3× better text, and from unusable to second place on tables. Any comparison's numbers hinge on serving decisions that almost nobody documents.

5. The quality-price frontier is on-prem. No cloud API dominates naktor-extractor on quality-per-euro. And there is a second axis price doesn't capture: governance — everything sent to Mistral leaves your infrastructure, which for banking, healthcare or government can be a hard disqualifier before price is even discussed. naktor-extractor deploys on your infra, at $0.99/1,000, with top-table quality; and on documents with a native text layer (most of the real world) the cost drops much further — something this 100%-scanned corpus doesn't exercise.

6. (Post-benchmark) The benchmark audited us back. Analyzing page by page where we lost, the whole text gap concentrated in ~5 pages whose table titles, figure captions and sub-questions were missing from our markdown: they were extracted fine, but our rendering dropped them. Noise-free A/B re-render (same outputs, only the rendering changes): text 0.088 → 0.064 (Chinese −36%), reading order 0.176 → 0.152. The biggest return on building a reproducible benchmark was being able to audit our own product with it; the fix is committed with its tests.

Limitations (read before quoting us)

  • 100 pages. Enough for large separations (15 TEDS points), fragile for fine differences: averages over small subsets (14 formula pages) move ±0.05 on a single page.
  • 100% scanned corpus (OmniDocBench ships images): on documents with a native text layer (most of the real world) the cost drops much further — something this corpus doesn't measure —, and classic CPU tools compete on their worst terrain.
  • Deliberately hard-skewed sample: do not compare these absolutes with the official leaderboard (sanity check: on the English pages it scores 0.036 vs the published 0.033 — the harness is wired correctly).
  • No formula CDM (edit distance only) and no Azure or AWS Textract (no credentials during execution; to be added in a revision).
  • Marker was measured but excluded from the main table: our setup (sequential local CPU) yields a cost 100× outside its representative range. Publishing it as «Marker» would mislead; it will be re-evaluated in GPU mode.
  • LLM judges are non-deterministic (±2 points between runs) and carry known biases; hence they are the secondary layer, and there are two from different families.
  • The authors sell naktor-extractor. The antidote is reproducibility: for the public systems, the harness, page IDs, seeds and configs are published; anyone can re-derive their rows.

Reproducibility

Dataset: OmniDocBench revision aa1ee96 (HuggingFace, free download). We publish: the seeded sampling script (corpus/sample.py, seed 20260703), the 100 page IDs, the runners for the public systems with exact configs, the official-harness driver, judge prompts, and this analysis. We do not republish the images or the extracted markdown (corpus copyright) — they regenerate from the scripts in ~1h and ~$15.

Total spend for this benchmark: ~$40 (cloud APIs, LLM judges and compute). Dismantling a sector's marketing costs less than a nice dinner.

We extract documents for systems that can't afford a wrong digit

Questions, rebuttals, or a system you want added to the table? Let's talk.

See naktor-extractor