CodeSOTA · Research graphPublic measurement and evidence layer · August 2026

Question-centric research infrastructure

The frontier is not a table. It is a chain of evidence.

CodeSOTA now connects benchmark measurements to the questions they answer, the experiments that produced them, and stable claims anyone can inspect or reproduce.

7
public research questions
18
testable hypotheses, including refuted branches
14
inspectable experiment protocols
32
evidence attachments beneath 18 public claims
6
cumulative findings and next steps

Six objects carry the research.

A leaderboard, a paper figure and a decision are projections of the same graph—not separate sources of truth.

Question

what we need to learn

Hypothesis

a testable answer

Experiment

the design and runs

Evidence

results with provenance

Claim

what the evidence supports

New question

what the result opens

Start with a research programme that already ran.

Parameter Golf becomes the first imported Slayer graph: source-pinned leaderboard rows, experiment protocols, typed lineage and evidence-backed claims.

Dataset 001

OpenAI Parameter Golf

How do we minimize language-model BPB under a 16 MB artifact budget?

48
accepted rows
1.0565
best pinned BPB
5
graph experiments
Open dataset →

Start with what remains unknown.

Each question preserves current belief, competing hypotheses, negative results and the next useful experiment.

RQ-0001

Which evaluations provide useful signal for 25M–500M language models?

activeweak evidence

Training loss clearly moves at 25M parameters, but comparable downstream evidence across the 25M–500M range is still missing.

RQ-0002

What does a single Document AI composite score hide?

partially answeredstrong evidence

The verified Mistral OCR 3 and ClearOCR results show that aggregate ranking is not a substitute for inspecting component metrics.

RQ-0003

When is an OCR evaluation large enough to compare?

openmoderate evidence

Existing runs expose the comparability problem; a controlled repeated protocol is still needed to estimate sample stability.

RQ-0004

Which MMLU results are actually comparable?

activemoderate evidence

Protocol heterogeneity is directly visible in current CodeSOTA evidence; a normalized comparison set remains open work.

RQ-0005

When should a benchmark move from leading to primary to control?

openweak evidence

Pythia reference results suggest overlapping dynamic ranges, but the promotion thresholds need repeated measurements and confidence intervals.

RQ-0006

Which reasoning benchmarks still separate frontier models?

partially answeredmoderate evidence

In this matched snapshot ARC-AGI-1 spans 31.4 points, ARC-Challenge spans 0.8, and GSM8K spans 0. Protocol-matched reruns are still required before calling this an intrinsic benchmark property.

RQ-PG-0001

How do we minimize language-model BPB under a 16 MB artifact budget?

partially answeredstrong evidence

The pinned leaderboard moves from the 1.2244 baseline to 1.0565 BPB. The selected #1394–#1530 milestones improve from 1.08563 to 1.07336, but they are a frontier progression rather than one literal code lineage.

A score is a projection. Evidence is the source of truth.

Every statement below has a permanent URL and an evidence ledger that resolves back to registry rows or external runs.

C-0001

The BDH-25M-PL run reduced validation loss from an approximately 5.6 random baseline to 1.41 after 10,000 steps on 100M Polish byte tokens.

source verified1 evidence record
C-0002

BDH-25M-PL has not been evaluated on ARC-AGI; Pathway's reported ARC-AGI result belongs to a separate 150M BDH-CQ system.

source verified1 evidence record
C-0003

CodeSOTA independently evaluated Mistral OCR 3 at 79.75 composite on the full 1,355-image OmniDocBench run.

CodeSOTA verified1 evidence record
C-0004

Mistral OCR 3's verified OmniDocBench component results differ materially across text, tables, formulae and reading order.

CodeSOTA verified4 evidence records
C-0005

ClearOCR retains strong verified text recognition while its verified full-document composite is 31.7.

CodeSOTA verified2 evidence records
C-0006

ClearOCR's verified OmniDocBench table TEDS is 0.8 because the pipeline outputs tables as plain text rather than structured tables.

CodeSOTA verified1 evidence record
C-0007

The stored OCRBench-EN runs shown here are not direct reproductions because their model, sample count and valid-sample coverage differ.

source verified3 evidence records
C-0008

A run-level OCR metric is incomplete evidence unless total and valid sample counts remain inspectable.

source verified1 evidence record
C-0009

CodeSOTA's MMLU evidence includes zero-shot chain-of-thought, five-shot and Global MMLU Lite rows, so the benchmark name alone does not establish comparability.

source verified3 evidence records
C-0010

The published CodeSOTA ladder shows SciQ separating earlier than ARC-Challenge across the Pythia 70M–6.9B reference series.

source verified1 evidence record
C-0011

ARC-AGI-1 spans 31.4 percentage points across the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.

CodeSOTA verified3 evidence records
C-0012

ARC-Challenge spans only 0.8 percentage points across the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.

CodeSOTA verified3 evidence records
C-0013

GSM8K reports 99.0 for all three models in the matched o3, o4-mini and Gemini 2.5 Pro CodeSOTA snapshot.

CodeSOTA verified3 evidence records
C-PG-1394

Merged Parameter Golf PR #1394 reports a 5-seed FineWeb validation result of 1.08563 BPB under the challenge constraints.

source verified1 evidence record
C-PG-1413

Merged Parameter Golf PR #1413 reports a 3-seed FineWeb validation result of 1.08279 BPB under the challenge constraints.

source verified1 evidence record
C-PG-1493

Merged Parameter Golf PR #1493 reports a 3-seed FineWeb validation result of 1.081 BPB under the challenge constraints.

source verified1 evidence record
C-PG-1514

Merged Parameter Golf PR #1514 reports a 3-seed FineWeb validation result of 1.07983 BPB under the challenge constraints.

source verified1 evidence record
C-PG-1530

Merged Parameter Golf PR #1530 reports a 3-seed FineWeb validation result of 1.07336 BPB under the challenge constraints.

source verified1 evidence record

The smallest useful loop ends by changing what we know.

Forks and reproductions begin as local, portable protocol manifests. Publishing is a deliberate final step.

01Read a claim
02Inspect evidence
03Fork protocol
04Run a variant
05Add evidence
06Change the graph