AI2 Reasoning Challenge.

7,787 science questions requiring reasoning. Challenge set contains harder questions that retrieval fails on.

Paper ↗Download dataset Submit a result ↵

§ 01 · Leaderboard

Best published scores.

10 results indexed across 1 metric. Shaded row marks current SOTA; ties broken by submission date.

Primary: accuracy · higher is better

accuracy· primary

10 rows

#	Model	Org	Submitted	Paper / code	accuracy
01	o3API	OpenAI	Mar 2026	openai-simple-evals	98.10
02	Gemini 2.5 ProAPI	Google	Mar 2026	google-technical-report	97.80
03	Llama-4-MaverickOSS	Meta	Mar 2026	meta-blog	97.40
04	o4-miniAPI	OpenAI	Mar 2026	openai-simple-evals	97.30
05	DeepSeek R1OSS	DeepSeek	Mar 2026	arxiv	97.10
06	Llama 3.1 405BOSS	Meta	Mar 2026	meta-modelcard	96.90
07	Claude 3.5 SonnetAPI	Anthropic	Dec 2025	anthropic-blog	96.70
08	GPT-4oAPI	OpenAI	Dec 2025	openai-blog	96.40
09	Gemini 1.5 ProAPI	Google	Dec 2025	google-blog	94.80
10	Llama 3 70BOSS	Meta	Dec 2025	meta-blog	93

Fig 2 · Rows sorted by score within each metric. Shaded row marks SOTA. Dates reflect model or paper release where available, otherwise the date Codesota accessed the source.

§ 03 · Progress

2 steps
of state of the art.

Each row below marks a model that broke the previous record on accuracy. Intermediate submissions are kept in the leaderboard above; only SOTA-setting entries are re-listed here.

Higher scores win. Each subsequent entry improved upon the previous best.

SOTA line · accuracy

Dec 17, 2025Claude 3.5 SonnetAnthropic96.70
Mar 27, 2026o3OpenAI98.10

Fig 3 · SOTA-setting models only. 2 entries span Dec 2025 → Mar 2026.

§ 06 · Contribute

Have a score that beats
this table?

Submit a checkpoint and a reproduction script. We will run it, publish the score, and — if it takes the top — annotate the step on the progress chart with your name.

Submit a result ↵Read submission guide

What a submission needs

01A public checkpoint or API endpoint
02A reproduction script with frozen commit + seed
03Declared evaluation environment (Python, deps)
04One row per metric declared by this dataset
05A contact so we can follow up on discrepancies

AI2 Reasoning Challenge.

Best published scores.

2 stepsof state of the art.

Neighbouring benchmarks.

Have a score that beatsthis table?

2 steps
of state of the art.

Have a score that beats
this table?