CodeSOTA · Research question · RQ-0003Updated 2026-08-20
What we want to learn

When is an OCR evaluation large enough to compare?

How much do sample count, valid-sample filtering and run configuration change the apparent quality of an OCR system?

Status
open
Evidence
moderate evidence
Graph
2 hypotheses · 1 experiments
Claims
2 public · 4 attachments
Reproductions
0 recorded

Existing runs expose the comparability problem; a controlled repeated protocol is still needed to estimate sample stability.

Why this question exists.

CodeSOTA has execution-level OCR runs ranging from 100 to 1,000 samples. They are evidence, but they are not automatically comparable reproductions.

Why it matters. Treating differently scoped runs as direct replications creates false certainty and can reverse model decisions.

Keep the supported and refuted branches.

A negative result remains visible research memory. Each branch links its protocol and the claim the evidence supports or refutes.

H-0006testing

Sample count changes apparent OCR quality

OCR quality estimates from a 100-sample subset are not stable enough to stand in for a 1,000-sample evaluation.

H-0007supported

Valid-sample filtering needs provenance

The valid/total sample ratio is necessary evidence for interpreting OCR run metrics.

What the graph can currently say.

F-0003 · investigating

Run comparability is a protocol property

Sharing a benchmark label is insufficient for reproduction. Sample scope, validity filtering and configuration must match or be recorded as deviations.

Visible ownership.

Stable ID
RQ-0003
Created by
Kacper Wikiel
Contributors
Kacper Wikiel · CodeSOTA evaluation programme
Visibility
public