CodeSOTA · Experiment · E-0008Portable protocol · public evidence graph
Why this run exists

ARC-Challenge matched-model audit

Measure the visible score range for the same three frontier models on ARC-Challenge.

Experiment
E-0008
Status
completed
Outcome
refuted
Forks
2 public or local starts recorded
Owner
Kacper Wikiel

Method before metrics.

Resolve the three verified benchmark_result rows live and compute max minus min. Source notes remain visible because the Gemini row is zero-shot CoT while OpenAI rows are described as zero-shot.

Hypothesis
H-0012 · ARC-Challenge produces at least a five-point range across o3, o4-mini and Gemini 2.5 Pro in the verified snapshot.

The comparison contract.

A reproduction matches this protocol. A fork changes it and declares the deviation.

dataset
arc-challenge
controlled Variables
model set · metric
independent Variable
model
benchmarks
accuracy
screening threshold points
5

What this experiment produced or evaluated.

Uses
benchmark:arc-challenge

Change the evidence, not just the discussion.

Start locally, inspect the portable manifest, and sign in only if you choose to publish.

Attach a run to this protocol.

Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.

Stable, portable, attributable.

Stable ID
E-0008
Visibility
public
Created
2026-08-20
Updated
2026-08-20