What does a single Document AI composite score hide?
Can one composite score represent text, tables, formulae and reading order without hiding practically important failure modes?
Why this question exists.
OmniDocBench compresses heterogeneous document behaviours into a headline number. CodeSOTA also stores the component metrics and independently evaluated results.
Why it matters. Buyers and researchers can choose the wrong system when a respectable aggregate conceals a near-total failure on a required document structure.
Keep the supported and refuted branches.
A negative result remains visible research memory. Each branch links its protocol and the claim the evidence supports or refutes.
Composite scores mask structure failures
A single document parsing composite can hide practically decisive table or formula failures.
Component metrics change OCR decisions
Inspecting text, table, formula and reading-order metrics changes the preferred OCR system for task-specific deployments.
What the graph can currently say.
CodeSOTA independently evaluated Mistral OCR 3 at 79.75 composite on the full 1,355-image OmniDocBench run.
CodeSOTA verified1 evidence recordC-0004Mistral OCR 3's verified OmniDocBench component results differ materially across text, tables, formulae and reading order.
CodeSOTA verified4 evidence recordsC-0005ClearOCR retains strong verified text recognition while its verified full-document composite is 31.7.
CodeSOTA verified2 evidence recordsC-0006ClearOCR's verified OmniDocBench table TEDS is 0.8 because the pipeline outputs tables as plain text rather than structured tables.
CodeSOTA verified1 evidence recordDocument composites need component evidence
Two independently evaluated systems can be useful for different jobs even when one aggregate score is much lower; the component metrics explain why.
Visible ownership.
- Stable ID
- RQ-0002
- Created by
- Kacper Wikiel
- Contributors
- Kacper Wikiel · CodeSOTA evaluation programme
- Visibility
- public