CodeSOTA · Experiment · E-0005Portable protocol · public evidence graph
OCRBench-EN sample-scope audit
Determine which stored OCR runs can be compared and which need a controlled reproduction.
01 · Design
Method before metrics.
Inspect total samples, valid samples, latency and error metrics before treating any two runs as direct evidence of model differences.
- Hypothesis
- H-0006 · OCR quality estimates from a 100-sample subset are not stable enough to stand in for a 1,000-sample evaluation.
- Hypothesis
- H-0007 · The valid/total sample ratio is necessary evidence for interpreting OCR run metrics.
02 · Protocol
The comparison contract.
A reproduction matches this protocol. A fork changes it and declares the deviation.
- dataset
- ocrbench-en
- independent Variable
- model and sample scope
- benchmarks
- accuracy · exact accuracy · CER · WER · latency
03 · Evidence outputs
What this experiment produced or evaluated.
- Produced
- Tesseract 5 · OCRBench-EN · 100 samples
external-run:ER-0002 - Produced
- dots.ocr · OCRBench-EN · 100 samples
external-run:ER-0003 - Produced
- PaddleOCR-VL · OCRBench-EN · 1,000 samples
external-run:ER-0004
Contribute
Change the evidence, not just the discussion.
Start locally, inspect the portable manifest, and sign in only if you choose to publish.
Add evidence
Attach a run to this protocol.
Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.
04 · Record
Stable, portable, attributable.
- Stable ID
- E-0005
- Visibility
- public
- Created
- 2026-08-20
- Updated
- 2026-08-20