When is an OCR evaluation large enough to compare?
How much do sample count, valid-sample filtering and run configuration change the apparent quality of an OCR system?
Why this question exists.
CodeSOTA has execution-level OCR runs ranging from 100 to 1,000 samples. They are evidence, but they are not automatically comparable reproductions.
Why it matters. Treating differently scoped runs as direct replications creates false certainty and can reverse model decisions.
Keep the supported and refuted branches.
A negative result remains visible research memory. Each branch links its protocol and the claim the evidence supports or refutes.
Sample count changes apparent OCR quality
OCR quality estimates from a 100-sample subset are not stable enough to stand in for a 1,000-sample evaluation.
Valid-sample filtering needs provenance
The valid/total sample ratio is necessary evidence for interpreting OCR run metrics.
What the graph can currently say.
The stored OCRBench-EN runs shown here are not direct reproductions because their model, sample count and valid-sample coverage differ.
source verified3 evidence recordsC-0008A run-level OCR metric is incomplete evidence unless total and valid sample counts remain inspectable.
source verified1 evidence recordRun comparability is a protocol property
Sharing a benchmark label is insufficient for reproduction. Sample scope, validity filtering and configuration must match or be recorded as deviations.
Visible ownership.
- Stable ID
- RQ-0003
- Created by
- Kacper Wikiel
- Contributors
- Kacper Wikiel · CodeSOTA evaluation programme
- Visibility
- public