CodeSOTA · Experiment · E-0002Portable protocol · public evidence graph
Why this run exists

Controlled 25M / 50M / 150M / 350M benchmark ladder

Measure where BPB, BLiMP, SciQ, ARC-Easy, LAMBADA and MMLU begin separating otherwise comparable base models.

Experiment
E-0002
Status
planned
Outcome
pending
Forks
8 public or local starts recorded
Owner
Kacper Wikiel

Method before metrics.

Train a fixed model family at four scales, preserve the data/token ratio, and evaluate each checkpoint with the same harness and multiple seeds.

Hypothesis
H-0002 · SciQ provides above-chance separation at smaller parameter scales than standard MMLU.
Hypothesis
H-0003 · Standard MMLU accuracy remains too close to chance to be a primary KPI for base models below roughly 300M parameters.
Hypothesis
H-0010 · A useful benchmark progresses from leading indicator to primary KPI to control as model scale increases.

The comparison contract.

A reproduction matches this protocol. A fork changes it and declares the deviation.

model Family
decoder-only base language model
controlled Variables
tokenizer · dataset mixture · training token ratio · evaluation prompts
independent Variable
parameter count
seeds
42 · 43 · 44
benchmarks
BPB · BLiMP · SciQ · ARC-Easy · LAMBADA · MMLU

What this experiment produced or evaluated.

This planned experiment has not produced public evidence yet.

Change the evidence, not just the discussion.

Start locally, inspect the portable manifest, and sign in only if you choose to publish.

Attach a run to this protocol.

Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.

Stable, portable, attributable.

Stable ID
E-0002
Visibility
public
Created
2026-08-20
Updated
2026-08-20