what we need to learn
Question-centric research infrastructure
The frontier is not a table. It is a chain of evidence.
CodeSOTA now connects benchmark measurements to the questions they answer, the experiments that produced them, and stable claims anyone can inspect or reproduce.
- 7
- public research questions
- 18
- testable hypotheses, including refuted branches
- 14
- inspectable experiment protocols
- 32
- evidence attachments beneath 18 public claims
- 6
- cumulative findings and next steps
Six objects carry the research.
A leaderboard, a paper figure and a decision are projections of the same graph—not separate sources of truth.
a testable answer
the design and runs
results with provenance
what the evidence supports
what the result opens
Start with a research programme that already ran.
Parameter Golf becomes the first imported Slayer graph: source-pinned leaderboard rows, experiment protocols, typed lineage and evidence-backed claims.
OpenAI Parameter Golf
How do we minimize language-model BPB under a 16 MB artifact budget?
- 48
- accepted rows
- 1.0565
- best pinned BPB
- 5
- graph experiments
Start with what remains unknown.
Each question preserves current belief, competing hypotheses, negative results and the next useful experiment.
Which evaluations provide useful signal for 25M–500M language models?
activeweak evidenceTraining loss clearly moves at 25M parameters, but comparable downstream evidence across the 25M–500M range is still missing.
RQ-0002What does a single Document AI composite score hide?
partially answeredstrong evidenceThe verified Mistral OCR 3 and ClearOCR results show that aggregate ranking is not a substitute for inspecting component metrics.
RQ-0003When is an OCR evaluation large enough to compare?
openmoderate evidenceExisting runs expose the comparability problem; a controlled repeated protocol is still needed to estimate sample stability.
RQ-0004Which MMLU results are actually comparable?
activemoderate evidenceProtocol heterogeneity is directly visible in current CodeSOTA evidence; a normalized comparison set remains open work.
RQ-0005When should a benchmark move from leading to primary to control?
openweak evidencePythia reference results suggest overlapping dynamic ranges, but the promotion thresholds need repeated measurements and confidence intervals.
RQ-0006Which reasoning benchmarks still separate frontier models?
partially answeredmoderate evidenceIn this matched snapshot ARC-AGI-1 spans 31.4 points, ARC-Challenge spans 0.8, and GSM8K spans 0. Protocol-matched reruns are still required before calling this an intrinsic benchmark property.
RQ-PG-0001How do we minimize language-model BPB under a 16 MB artifact budget?
partially answeredstrong evidenceThe pinned leaderboard moves from the 1.2244 baseline to 1.0565 BPB. The selected #1394–#1530 milestones improve from 1.08563 to 1.07336, but they are a frontier progression rather than one literal code lineage.
A score is a projection. Evidence is the source of truth.
Every statement below has a permanent URL and an evidence ledger that resolves back to registry rows or external runs.
The BDH-25M-PL run reduced validation loss from an approximately 5.6 random baseline to 1.41 after 10,000 steps on 100M Polish byte tokens.
source verified1 evidence recordC-0002BDH-25M-PL has not been evaluated on ARC-AGI; Pathway's reported ARC-AGI result belongs to a separate 150M BDH-CQ system.
source verified1 evidence recordC-0003CodeSOTA independently evaluated Mistral OCR 3 at 79.75 composite on the full 1,355-image OmniDocBench run.
CodeSOTA verified1 evidence recordC-0004Mistral OCR 3's verified OmniDocBench component results differ materially across text, tables, formulae and reading order.
CodeSOTA verified4 evidence recordsC-0005ClearOCR retains strong verified text recognition while its verified full-document composite is 31.7.
CodeSOTA verified2 evidence recordsC-0006ClearOCR's verified OmniDocBench table TEDS is 0.8 because the pipeline outputs tables as plain text rather than structured tables.
CodeSOTA verified1 evidence recordC-0007The stored OCRBench-EN runs shown here are not direct reproductions because their model, sample count and valid-sample coverage differ.
source verified3 evidence recordsC-0008A run-level OCR metric is incomplete evidence unless total and valid sample counts remain inspectable.
source verified1 evidence recordC-0009CodeSOTA's MMLU evidence includes zero-shot chain-of-thought, five-shot and Global MMLU Lite rows, so the benchmark name alone does not establish comparability.
source verified3 evidence recordsC-0010The published CodeSOTA ladder shows SciQ separating earlier than ARC-Challenge across the Pythia 70M–6.9B reference series.
source verified1 evidence recordC-0011ARC-AGI-1 spans 31.4 percentage points across the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.
CodeSOTA verified3 evidence recordsC-0012ARC-Challenge spans only 0.8 percentage points across the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.
CodeSOTA verified3 evidence recordsC-0013GSM8K reports 99.0 for all three models in the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.
CodeSOTA verified3 evidence recordsC-PG-1394Merged Parameter Golf PR #1394 reports a 5-seed FineWeb validation result of 1.08563 BPB under the challenge constraints.
source verified1 evidence recordC-PG-1413Merged Parameter Golf PR #1413 reports a 3-seed FineWeb validation result of 1.08279 BPB under the challenge constraints.
source verified1 evidence recordC-PG-1493Merged Parameter Golf PR #1493 reports a 3-seed FineWeb validation result of 1.081 BPB under the challenge constraints.
source verified1 evidence recordC-PG-1514Merged Parameter Golf PR #1514 reports a 3-seed FineWeb validation result of 1.07983 BPB under the challenge constraints.
source verified1 evidence recordC-PG-1530Merged Parameter Golf PR #1530 reports a 3-seed FineWeb validation result of 1.07336 BPB under the challenge constraints.
source verified1 evidence recordThe smallest useful loop ends by changing what we know.
Forks and reproductions begin as local, portable protocol manifests. Publishing is a deliberate final step.