CodeSOTA · Research question · RQ-0002Updated 2026-08-20
What we want to learn

What does a single Document AI composite score hide?

Can one composite score represent text, tables, formulae and reading order without hiding practically important failure modes?

Status
partially answered
Evidence
strong evidence
Graph
2 hypotheses · 2 experiments
Claims
4 public · 8 attachments
Reproductions
0 recorded

The verified Mistral OCR 3 and ClearOCR results show that aggregate ranking is not a substitute for inspecting component metrics.

Why this question exists.

OmniDocBench compresses heterogeneous document behaviours into a headline number. CodeSOTA also stores the component metrics and independently evaluated results.

Why it matters. Buyers and researchers can choose the wrong system when a respectable aggregate conceals a near-total failure on a required document structure.

Keep the supported and refuted branches.

A negative result remains visible research memory. Each branch links its protocol and the claim the evidence supports or refutes.

What the graph can currently say.

F-0002 · confirmed

Document composites need component evidence

Two independently evaluated systems can be useful for different jobs even when one aggregate score is much lower; the component metrics explain why.

Visible ownership.

Stable ID
RQ-0002
Created by
Kacper Wikiel
Contributors
Kacper Wikiel · CodeSOTA evaluation programme
Visibility
public