CodeSOTA · Research question · RQ-0005Updated 2026-08-20
What we want to learn

When should a benchmark move from leading to primary to control?

At what model scale does each language benchmark begin providing useful signal, and when does it become too saturated to remain a primary KPI?

Status
open
Evidence
weak evidence
Graph
1 hypotheses · 1 experiments
Claims
1 public · 1 attachments
Reproductions
0 recorded

Pythia reference results suggest overlapping dynamic ranges, but the promotion thresholds need repeated measurements and confidence intervals.

Why this question exists.

The CodeSOTA benchmark ladder proposes lifecycle roles rather than one permanent suite for every parameter scale.

Why it matters. A benchmark should enter the decision set when it separates checkpoints and leave it before ceiling effects dominate.

Keep the supported and refuted branches.

A negative result remains visible research memory. Each branch links its protocol and the claim the evidence supports or refutes.

H-0010testing

Benchmark roles move with scale

A useful benchmark progresses from leading indicator to primary KPI to control as model scale increases.

What the graph can currently say.

Visible ownership.

Stable ID
RQ-0005
Created by
Kacper Wikiel
Contributors
Kacper Wikiel
Visibility
public