Which OCR model is best on average?
A single high score is easy to game — train on the test set, hand-tune one document type, publish a paper. Average performance across many benchmarks is harder to fake.
We rank every OCR model that placed on at least 2 of 9 public OCR benchmarks. Then — where we’ve verified the model ourselves — we show our own number next to the public consensus.
Within each benchmark we rank every model that has a score, then map the rank to a 0–100 percentile (top = 100). This neutralises that CER is lower-better while OmniDoc composite is higher-better — both end up on the same 0–100 axis.
Power score is the unweighted mean of a model's percentiles across the benchmarks where it has a score. We require a minimum of 2 benchmarks — one strong showing isn't enough.
Where CodeSOTA has run its own eval (currently 2 of 29 ranked models), the right-most column shows that score. When the public consensus and our number disagree, that disagreement is the most useful thing on this page.
The Power Ranking, 29 models.
Sorted by average percentile across the eight axes. Coverage column is load-bearing — a model on top with 2/8 is making a narrower claim than one on top with 6/8.
Pills below each model show per-benchmark percentile. Copper = top quartile (≥75), grey = middle, faded = bottom quartile.
| # | Model | Power | Coverage | CodeSOTA verified | Per-benchmark percentile |
|---|---|---|---|---|---|
| 01 | Qwen2.5-VL-72B | 97.0 | 2 / 9 | not yet | OCRBench EN 94OCRBench ZH 100 |
| 02 | PaddleOCR-VL | 81.3 | 3 / 9 | not yet | OmniDoc 90OmniDoc 85olmOCR 69 |
| 03 | Gemini 2.5 Pro | 80.0 | 5 / 9 | not yet | OmniDoc 54OCRBench EN 81OCRBench ZH 90MME-VideoOCR 100Thai-OCR 75 |
| 04 | Qianfan-OCR | 79.0 | 4 / 9 | not yet | OmniDoc 95OCRBench EN 75OCRBench ZH 80olmOCR 66 |
| 05 | PaddleOCR-VL-1.5 | 78.0 | 2 / 9 | not yet | OmniDoc 97olmOCR 59 |
| 06 | Intern-S1-Pro | 77.0 | 2 / 9 | not yet | OCRBench EN 84OCRBench ZH 70 |
| 07 | Ovis2.5-9B | 75.0 | 2 / 9 | not yet | OCRBench EN 100OCRBench ZH 50 |
| 08 | Falcon-OCR | 67.0 | 2 / 9 | not yet | OmniDoc 62olmOCR 72 |
| 09 | Claude Sonnet 4 | 65.5 | 2 / 9 | not yet | OCRBench EN 31Thai-OCR 100 |
| 10 | DeepSeek-OCR-2 | 61.5 | 2 / 9 | not yet | OmniDoc 82olmOCR 41 |
| 11 | Gemini 1.5 Pro | 60.0 | 2 / 9 | not yet | CC-OCR 100MME-VideoOCR 20 |
| 12 | minicpm-v-4.5-8b | 59.5 | 2 / 9 | not yet | OCRBench EN 59OCRBench ZH 60 |
| 13 | GLM-OCR | 58.0 | 2 / 9 | not yet | OmniDoc 100olmOCR 16 |
| 14 | dots.ocr 3B | 56.0 | 2 / 9 | not yet | OmniDoc 56olmOCR 56 |
| 15 | GPT-4o | 52.0 | 4 / 9 | not yet | OCRBench EN 72CC-OCR 25MME-VideoOCR 40KITAB 71 |
| 16 | sail-vl2-8b | 51.5 | 2 / 9 | not yet | OCRBench EN 63OCRBench ZH 40 |
| 17 | MonkeyOCR-pro-3B | 49.0 | 2 / 9 | not yet | OmniDoc 67olmOCR 31 |
| 18 | MinerU 2.5 | 47.7 | 3 / 9 | not yet | OmniDoc 77olmOCR 47olmOCR 19 |
| 19 | GPT-4o Mini | 45.5 | 2 / 9 | not yet | OCRBench EN 34KITAB 57 |
| 20 | claude-3.5-sonnet | 40.0 | 2 / 9 | not yet | OCRBench EN 50OCRBench ZH 30 |
| 21 | Qwen2.5-VL 72B | 40.0 | 2 / 9 | not yet | MME-VideoOCR 80Thai-OCR 0 |
| 22 | Mistral OCR 3 | 34.0 | 2 / 9 | 94.9 %Internal acc3.7 %CodeSOTA CER7.1 %CodeSOTA WER | OmniDoc 18olmOCR 50 |
| 23 | DeepSeek-OCR | 33.0 | 2 / 9 | not yet | OmniDoc 41olmOCR 25 |
| 24 | Qwen2-VL-72B | 33.0 | 2 / 9 | not yet | OCRBench EN 56OCRBench ZH 10 |
| 25 | InternVL2.5-78B | 30.5 | 2 / 9 | not yet | OCRBench EN 41OCRBench ZH 20 |
| 26 | gpt-4o-2024 | 26.5 | 2 / 9 | not yet | OCRBench EN 53OCRBench ZH 0 |
| 27 | Qwen2.5-VL 32B | 25.0 | 2 / 9 | not yet | MME-VideoOCR 0Thai-OCR 50 |
| 28 | olmOCR | 24.0 | 2 / 9 | not yet | OmniDoc 26olmOCR 22 |
| 29 | mistral-ocr-2512 | 9.0 | 2 / 9 | 1.22 p/spages/s | OmniDoc 15OCRBench EN 3 |
Public benchmarks aren’t enough.
Three problems compound. One: popular OCR benchmarks (OmniDoc, OCRBench, olmOCR) are easy to overfit — six months after a paper ships, the test set is in the next training run. Two: they miss the document types that actually pay rent — Polish invoices, German handwritten medical forms, scanned legacy PDFs with deliberate redactions. Three: a vendor’s self-reported score is a marketing artefact until somebody else runs the same eval.
Our verified column closes the third gap. The first two we close with a hold-out architecture: methodology and sample items are public, the actual test set rotates quarterly and stays private — so even when our questions eventually leak into a training corpus, they’re no longer the questions we’re using.
Currently 2 of 29 models on this page have a CodeSOTA-verified score. Expanding that coverage is the work.
Want a model verified against your docs?
If you’re evaluating OCR for production and a model on this list doesn’t have a CodeSOTA-verified score, tell us. We’ll prioritise what real practitioners are about to deploy over what arXiv published last week.