CodeSOTA · Experiment · E-0002Portable protocol · public evidence graph
Controlled 25M / 50M / 150M / 350M benchmark ladder
Measure where BPB, BLiMP, SciQ, ARC-Easy, LAMBADA and MMLU begin separating otherwise comparable base models.
01 · Design
Method before metrics.
Train a fixed model family at four scales, preserve the data/token ratio, and evaluate each checkpoint with the same harness and multiple seeds.
- Hypothesis
- H-0002 · SciQ provides above-chance separation at smaller parameter scales than standard MMLU.
- Hypothesis
- H-0003 · Standard MMLU accuracy remains too close to chance to be a primary KPI for base models below roughly 300M parameters.
- Hypothesis
- H-0010 · A useful benchmark progresses from leading indicator to primary KPI to control as model scale increases.
02 · Protocol
The comparison contract.
A reproduction matches this protocol. A fork changes it and declares the deviation.
- model Family
- decoder-only base language model
- controlled Variables
- tokenizer · dataset mixture · training token ratio · evaluation prompts
- independent Variable
- parameter count
- seeds
- 42 · 43 · 44
- benchmarks
- BPB · BLiMP · SciQ · ARC-Easy · LAMBADA · MMLU
03 · Evidence outputs
What this experiment produced or evaluated.
This planned experiment has not produced public evidence yet.
Contribute
Change the evidence, not just the discussion.
Start locally, inspect the portable manifest, and sign in only if you choose to publish.
Add evidence
Attach a run to this protocol.
Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.
04 · Record
Stable, portable, attributable.
- Stable ID
- E-0002
- Visibility
- public
- Created
- 2026-08-20
- Updated
- 2026-08-20