CodeSOTA · Experiment · E-0008Portable protocol · public evidence graph
ARC-Challenge matched-model audit
Measure the visible score range for the same three frontier models on ARC-Challenge.
01 · Design
Method before metrics.
Resolve the three verified benchmark_result rows live and compute max minus min. Source notes remain visible because the Gemini row is zero-shot CoT while OpenAI rows are described as zero-shot.
- Hypothesis
- H-0012 · ARC-Challenge produces at least a five-point range across o3, o4-mini and Gemini 2.5 Pro in the verified snapshot.
02 · Protocol
The comparison contract.
A reproduction matches this protocol. A fork changes it and declares the deviation.
- dataset
- arc-challenge
- controlled Variables
- model set · metric
- independent Variable
- model
- benchmarks
- accuracy
- screening threshold points
- 5
03 · Evidence outputs
What this experiment produced or evaluated.
- Uses
- benchmark:arc-challenge
Contribute
Change the evidence, not just the discussion.
Start locally, inspect the portable manifest, and sign in only if you choose to publish.
Add evidence
Attach a run to this protocol.
Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.
04 · Record
Stable, portable, attributable.
- Stable ID
- E-0008
- Visibility
- public
- Created
- 2026-08-20
- Updated
- 2026-08-20