Codesota · Large Language ModelsThe frontier leaderboard, dated & sourcedIssue: April 22, 2026
Live registry · 7 benchmarks · 131 models ·

Large language models,
measured honestly.

Frontier model performance across knowledge, reasoning, math, code, and sustained tool-use — every score dated, every source linked, every benchmark described in its own words. No vendor spin, no collapsed averages.

Shaded rows mark current state of the art. Descriptions in serif; scores in tabular mono; navigation in sans.

§ 01 · Dashboard

Current frontier, not historical coverage.

One row per current frontier model, one column per benchmark. Cells show the published score; a dash means we have no verified result for that pair. GPT-4o and other older systems remain in benchmark histories, but not in this frontier shortlist.


Frontier rows
16
Benchmarks
7
Results
317
Last update
May 26, 2026
Current frontier models · tracked evidence
Copper cells mark benchmark leaders
#ModelVendorMMLUGPQA DiamondAIME 2025LiveCodeBench ProLiveCodeBench (classic)Tau2-BenchHumanity's Last Exam (HLE)Cov.
01Gemini 3.1 ProGoogle288746.44%2/7
02Gemini 3 FlashGoogle89.6%90.4%90.8%3/7
03DeepSeek V3.5DeepSeek88.2%1/7
04Mistral Large 3Mistral87.1%1/7
05MiniMax M2.5MiniMax86.5%1/7
06Claude Opus 4.6Anthropic91.2%91.3%34.44%3/7
07Gemini 3 Pro PreviewGoogle91.7%37.52%2/7
08Gemini 3 ProGoogle91.4%91.9%243969%38.3%5/7
09GPT-5OpenAI90.8%89%217685%25.32%5/7
10Grok 4xAI86.6%88%79%24.5%4/7
11Claude Opus 4.5Anthropic91.8%74.9%80%79%25.2%5/7
12GPT-5.2OpenAI92.4%73%27.8%3/7
13DeepSeek-R1-0528DeepSeek73.3%1/7
14o4-miniOpenAI90%77.6%92.7%209272.8%18.08%6/7
15Qwen3-235B-A22BAlibaba87.81%71.1%81.5%167370.7%5/7
16Kimi K2.5Moonshot AI86%24.37%2/7
Fig 2 · This is a current frontier shortlist, not a raw coverage ranking. GPT-4o, GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro and similar historical rows stay in the individual benchmark tables below for lineage, but are intentionally excluded here.
§ 02 · Coverage

7 benchmarks. Five axes.

Knowledge, reasoning, math, code, tools, and frontier difficulty. Each tile names the benchmark, the current SOTA score, and the leading model.

MMLU
92.9%64 models
o3
Knowledge
GPQA Diamond
91.9%74 models
Gemini 3 Pro
Knowledge
AIME 2025
99.9%22 models
Step-3.5-Flash PaCoRe
Math & Reasoning
LiveCodeBench Pro
2887Elo10 models
Gemini 3.1 Pro
Coding
LiveCodeBench (classic)
93.5%54 models
DeepSeek-V4-Pro Max
Coding
Tau2-Bench
89.7%19 models
GLM-5
Agentic & Tools
Humanity's Last Exam (HLE)
54%74 models
Kimi K2.6
Frontier Difficulty
Fig 3 · Current SOTA per benchmark with count of recorded submissions. Click any tile to scroll to its full leaderboard below.
§ 03 · Caveat

Read the numbers with care.

Benchmarks are instruments. Like all instruments they can be mis-calibrated, miscounted, or gamed.

A 2026 Berkeley RDI study found that eight major agent benchmarks — including SWE-bench Verified, Terminal-Bench, WebArena, OSWorld, GAIA, and FieldWorkArena — could be exploited to near-perfect scores without solving any tasks.

Failure modes included leaked reference answers, unsanitized eval(), prompt-injectable LLM judges, and scoring functions that skip correctness checks entirely. A 10-line conftest.py was enough to make every SWE-bench test report as passing.

Treat leaderboard position as a signal, not proof of capability — especially on agentic benchmarks where the evaluation environment is itself part of the attack surface. Held-out, contamination-resistant evals like HLE and LiveCodeBench Pro are more resistant, but not immune.

Read the full Berkeley RDI analysis →
§ 04 · Knowledge

Knowledge.

Breadth across 57+ subjects, graduate-level and multiple-choice.

The original MMLU: 15,908 four-choice questions across 57 subjects from elementary to professional level. Largely saturated at the frontier — top models cluster above 90%. For a harder variant see MMLU-Pro.

#ModelVendoraccuracy
01o3OpenAI92.9%
02GPT-5.2OpenAI92.4%
03Claude Opus 4.5Anthropic91.8%
04o1OpenAI91.8%
05Claude Opus 4.5Anthropic91.6%
06Gemini 3 ProGoogle91.4%
07Claude Opus 4.6Anthropic91.2%
08o1-previewOpenAI90.8%
09GPT-5OpenAI90.8%
10DeepSeek R1DeepSeek90.8%
11GPT-4.5 PreviewOpenAI90.8%
12Claude Sonnet 4.5Anthropic90.4%
13GPT-4.1OpenAI90.2%
14Claude Sonnet 4Anthropic90.1%
15o4-miniOpenAI90%
16GLM-4.5Zhipu AI90%
17Gemini 2.5 ProGoogle89.8%
18Gemini 3 FlashGoogle89.6%
19Llama 4 MaverickMeta89.4%
20Claude Opus 4Anthropic88.8%
21Qwen 3 72BAlibaba88.7%
22Llama 3.1 405BMeta88.6%
23DeepSeek-V3DeepSeek88.5%
24MiniMax-Text-01MiniMax88.5%
25Claude 3.5 SonnetAnthropic88.3%
26DeepSeek V3.5DeepSeek88.2%
27Qwen3-235B-A22BAlibaba87.81%
28Llama 4 405BMeta87.8%
29Qwen3-Coder-NextQwen87.73%
30Grok 2xAI87.5%
31Llama 3 (405B, Instruct)Meta87.3%
32Trinity Large PreviewArcee AI87.21%
33GPT-4oOpenAI87.2%
34Mistral Large 3Mistral87.1%
35LongCat-Flash-Omni86.81%
36Claude 3 OpusAnthropic86.8%
37GPT-4 TurboOpenAI86.7%
38Grok 4xAI86.6%
39MiniMax M2.5MiniMax86.5%
40Qwen2.5-72B-InstructAlibaba86.1%
41Kimi K2.5Moonshot AI86%
42o3-miniOpenAI85.9%
43Gemini 1.5 ProGoogle85.9%
44Step-3.5-Flash Base85.8%
45o1-miniOpenAI85.2%
46Qwen 3 14BAlibaba84.3%
47Phi-4 14BMicrosoft83.9%
48GPT-4o miniOpenAI82%
49Llama 3.1 70BMeta82%
50Qwen3-Omni-30B-A3B-Base-20250781.69%
51Qwen3-VL-8B-InstructQwen80.7%
52MiniCPM-o 4.5-Instruct77%
53Aria73.3%
54Apertus-70B-Instruct69.6%
55Llama 2 70B (5-shot)68.9%
56Chameleon 34B65.8%
57Apertus-70B65.2%
58LLaMA-65B63.4%
59OLMo-2-7B-1124 (olmOCR-peS2o)61.1%
60HRM-Text-1B60.7%
61BLT-Entropy 8B57.4%
62Helium54.3%
63BitNet b1.58 2B4T53.17%
64MoshiKyutai49.7%
Source: hendrycks/test (MMLU) · Saturated benchmark. Small score deltas at the top (90–93%) are within noise; treat rankings as a cluster, not a strict order.

198 expert-authored graduate-level questions in biology, chemistry, and physics. PhD-level specialists score ~65% on their own field. Designed to be impossible to Google.

#ModelVendoraccuracy
01Gemini 3 ProGoogle91.9%
02Claude Opus 4.6Anthropic91.3%
03Kimi K2.690.5%
04Gemini 3 FlashGoogle90.4%
05DeepSeek-V4-Pro MaxDeepSeek90.1%
06Claude Sonnet 4.6Anthropic89.9%
07GPT-5OpenAI89%
08Qwen3.5-397B-A17BAlibaba88.4%
09DeepSeek-V4-Flash MaxDeepSeek88.1%
10Grok 4xAI88%
11Qwen3.6-27B87.8%
12Kimi-K2.5Moonshot.AI87.6%
13Qwen3.5-122B-A10BAlibaba86.6%
14Gemini 2.5 Pro86.4%
15GLM-5.186.2%
16GLM-5Zhipu AI86%
17Qwen3.6-35B-A3B86%
18DeepSeek-V3.2-SpecialeDeepSeek85.7%
19GLM-4.7Zhipu AI85.7%
20Qwen3.5-27BAlibaba85.5%
21MiniMax-M2.5MiniMaxAI85.2%
22Step-3.5-Flash PaCoRe85%
23Gemma 4 31BGoogle84.3%
24Qwen3.5-35B-A3BAlibaba84.2%
25Gemini 2.5 ProGoogle84%
26Qwen3.5-Omni-Plus83.9%
27Step-3.5-Flash83.5%
28Gemini 2.5 FlashGoogle82.8%
29Gemini 2.5 Flash82.8%
30o3OpenAI82.8%
31DeepSeek-V3.2DeepSeek82.4%
32NVIDIA-Nemotron-3-Super-120B-A12B-BF1679.23%
33GLM-4.5Zhipu AI79.1%
34o4-miniOpenAI77.6%
35Qwen3-VL-235B-A22B-ThinkingQwen77.1%
36Claude Opus 4Anthropic76.7%
37o1OpenAI75.7%
38GLM-4.5-AirZhipu AI75%
39o3-miniOpenAI74.9%
40Claude Opus 4.5Anthropic74.9%
41Qwen3-Coder-NextQwen74.49%
42Qwen3-VL-235B-A22B-InstructQwen74.3%
43o1-previewOpenAI73.3%
44Qwen3-Omni-Flash-Thinking73.1%
45NVIDIA-Nemotron-3-Nano-30B-A3B-BF1673%
46DeepSeek R1DeepSeek71.5%
47Qwen3-235B-A22BAlibaba71.1%
48Qwen3-235B-A22BAlibaba71.1%
49ZAYA1-8BZ.ai71%
50Claude Sonnet 4Anthropic70%
51Llama 4 MaverickMeta69.8%
52GPT-4.5 PreviewOpenAI69.5%
53MiMo-V2.5-Pro66.7%
54GPT-4.1 miniOpenAI66.4%
55GPT-4.1OpenAI66.3%
56Trinity Large PreviewArcee AI63.32%
57o1-miniOpenAI60%
58Claude 3.5 SonnetAnthropic59.4%
59Grok 2xAI56%
60MiniMax-Text-01MiniMax54.4%
61Llama 3 (405B, Instruct)Meta51.1%
62Llama 3.1 405BMeta50.7%
63Claude 3 OpusAnthropic50.4%
64GPT-4oOpenAI49.9%
65Qwen2.5-Plus49.7%
66GPT-4 TurboOpenAI49.3%
67Qwen2.5-72B-InstructAlibaba49%
68Qwen2.5-VL-72B49%
69Gemini 1.5 ProGoogle46.2%
70Gemma 3 (27B, IT)42.4%
71Llama 3.1 70BMeta41.7%
72Step-3.5-Flash Base41.7%
73GPT-4o miniOpenAI40.2%
74Qwen3-VL-8B-InstructQwen34.7%
Source: arXiv:2311.12022 · Human expert baseline (non-specialist): 34%. PhD specialist: ~65%.
§ 05 · Math / Reasoning

Math & reasoning.

Olympiad-style short answer, released after model training cutoffs.

The 2025 American Invitational Mathematics Examination: 30 olympiad-style short-answer problems drawn after most 2024-era model training cutoffs. A primary frontier-math signal in recent reasoning-model reports.

#ModelVendoraccuracy
01Step-3.5-Flash PaCoRe99.9%
02Step-3.5-Flash97.3%
03Kimi-K2.5Moonshot.AI96.1%
04DeepSeek-V3.2-SpecialeDeepSeek96%
05SU-0194.6%
06Intern-S1-ProShanghai AI Lab93.1%
07DeepSeek-V3.2DeepSeek93.1%
08o4-miniOpenAI92.7%
09Qwen3-VL-235B-A22B-ThinkingQwen89.7%
10NVIDIA-Nemotron-3-Nano-30B-A3B-BF1689.1%
11Gemini 2.5 Pro88%
12o3OpenAI86.7%
13Gemini 2.5 ProGoogle86.7%
14Qwen3-Coder-NextQwen83.07%
15Qwen3-235B-A22BAlibaba81.5%
16Claude Opus 4.5Anthropic80%
17Qwen3-VL-235B-A22B-InstructQwen74.7%
18Qwen3-Omni-Flash-Thinking74%
19DeepSeek R1DeepSeek72%
20Gemini 2.5 Flash72%
21Qwen3-VL-8B-InstructQwen45.9%
22Trinity Large PreviewArcee AI24.36%
Source: maa.org/aime · Small test set (30 problems) — a single swing is ~3.3%. Numbers below are pass@1 unless otherwise noted.
§ 06 · Coding

Code.

Contest-style programming. Elo-rated or pass@1 on held-out problems.

The 2026 Elo-rated successor to classic LCB. Built by Olympiad medalists from continuously-updated Codeforces, ICPC and IOI problems. Each LLM is treated as a virtual Codeforces contestant and fit to a Bayesian MAP Elo on the standard Codeforces scale (~800 novice to ~3800 top human).

#ModelVendorElo
01Gemini 3.1 ProGoogle2887
02Gemini 3 ProGoogle2439
03GPT-5OpenAI2176
04o4-miniOpenAI2092
05Gemini 2.5 ProGoogle1769
06Qwen3-235B-A22BAlibaba1673
07Claude Sonnet 4.5Anthropic1412
08Gemini 2.5 FlashGoogle1288
09DeepSeek R1DeepSeek1161
10o3OpenAI1010
Source: livecodebenchpro.com · Elo rating comparable to the Codeforces human scale. Top human contestants sit around 3800; the strongest model on the board is Gemini 3 Pro at 2439.

Classic pass@1 LiveCodeBench — continuously updated with new contest problems from LeetCode, Codeforces, and AtCoder. Largely superseded by LCB Pro for frontier models, but preserved here for historical comparison across older models.

#ModelVendorpass-1
01DeepSeek-V4-Pro MaxDeepSeek93.5%
02Gemini 3 Pro PreviewGoogle91.7%
03DeepSeek-V4-Flash MaxDeepSeek91.6%
04Gemini 3 FlashGoogle90.8%
05Kimi K2.689.6%
06DeepSeek-V3.2-SpecialeDeepSeek88.7%
07Kimi-K2.5Moonshot.AI85%
08GPT-5OpenAI85%
09Qwen3.6-27B83.9%
10Qwen3.5-397B-A17BAlibaba83.6%
11DeepSeek-V3.2DeepSeek83.3%
12NVIDIA-Nemotron-3-Super-120B-A12B-BF1681.19%
13Qwen3.6-35B-A3B80.4%
14Gemma 4 31BGoogle80%
15Grok 4xAI79%
16Gemini 2.5 ProGoogle75.6%
17Intern-S1-ProShanghai AI Lab74.3%
18Gemini 2.5 Pro74.2%
19DeepSeek-R1-0528DeepSeek73.3%
20GLM-4.5Zhipu AI72.9%
21o4-miniOpenAI72.8%
22Qwen3-235B-A22BAlibaba70.7%
23GLM-4.5-AirZhipu AI70.7%
24Qwen3-235B-A22BAlibaba70.7%
25Qwen3-VL-235B-A22B-ThinkingQwen70.1%
26NVIDIA-Nemotron-3-Nano-30B-A3B-BF1668.3%
27o3-miniOpenAI66.9%
28DeepSeek R1DeepSeek65.9%
29o3OpenAI65.3%
30DeepSeek-R1-Distill-Llama-70BDeepSeek65.2%
31Gemini 2.5 FlashGoogle63.9%
32Kimi k1.5Moonshot AI62.5%
33DeepSeek-R1-Distill-Qwen-32BDeepSeek62.1%
34Gemini 2.5 Flash59.3%
35Qwen3-Coder-NextQwen58.93%
36Claude Opus 4Anthropic57.8%
37Qwen2.5-72B-Instruct55.5%
38GPT-4.1OpenAI54.4%
39Qwen3-VL-235B-A22B-InstructQwen54.3%
40Claude Sonnet 4Anthropic52.8%
41DeepSeek-v3-0324DeepSeek49.2%
42DeepSeek-V3DeepSeek49.2%
43GPT-4.1 miniOpenAI48.3%
44Qwen2.5-Coder 32BAlibaba47.8%
45Llama 4 MaverickMeta43.4%
46DeepSeek-Coder-V2-InstructDeepSeek43.4%
47GPT-4oOpenAI40.8%
48Qwen3-VL-8B-InstructQwen39.3%
49Gemma-3-27bGoogle39%
50Llama-4-ScoutMeta32.8%
51Gemma 3 12B ITGoogle DeepMind32%
52Gemma 3 (27B, IT)29.7%
53Codestral 22BMistral29.5%
54Gemma 3 4B ITGoogle DeepMind23%
Source: livecodebench.github.io · Problems released after model training cutoffs to prevent contamination.
§ 07 · Agentic / Tools

Agentic & tools.

Multi-turn tasks using real tools and databases; pass = full resolution.

Simulates real customer service interactions — agents use tools and databases to resolve tasks in retail and airline domains across multi-turn dialogues. Pass rate = task fully resolved.

#ModelVendoraccuracy
01GLM-5Zhipu AI89.7%
02Step-3.5-Flash88.2%
03Qwen3.5-397B-A17BAlibaba86.7%
04Qwen3.5-35B-A3BAlibaba81.2%
05Intern-S1-ProShanghai AI Lab80.9%
06DeepSeek-V3.2DeepSeek80.3%
07Qwen3.5-122B-A10BAlibaba79.5%
08Qwen3.5-27BAlibaba79%
09Claude Opus 4.5Anthropic79%
10Ling-2.6-1T78.36%
11SenseNova-U1-A3B-MoTSenseTime75.39%
12GPT-5.2OpenAI73%
13Gemini 3 ProGoogle69%
14Claude Sonnet 4.5Anthropic63%
15NVIDIA-Nemotron-3-Super-120B-A12B-BF1661.15%
16GPT-5.1OpenAI59%
17Gemini 2.5 ProGoogle54%
18Claude 3.7 SonnetAnthropic47%
19GPT-4oOpenAI36%
Source: sierra-research/tau2-bench · Average across 3 seeds per model.
§ 08 · Frontier

Frontier difficulty.

Designed to remain unsaturated for years. Even the leaders score low.

Humanity's Last Exam (HLE)

3,000 extremely hard questions across math, science, law, and humanities — contributed by domain experts worldwide. Designed to remain unsaturated for years. No tools allowed in this variant.

#ModelVendoraccuracy
01Kimi K2.654%
02MiMo-V2.5-Pro48%
03Gemini 3.1 ProGoogle46.44%
04GPT-5.4 ProOpenAI44.32%
05Muse SparkMeta40.56%
06Gemini 3 ProGoogle38.3%
07DeepSeek-V4-Pro MaxDeepSeek37.7%
08Gemini 3 Pro PreviewGoogle37.52%
09GPT-5.4OpenAI36.24%
10Claude Opus 4.7Anthropic36.2%
11DeepSeek-V4-Flash MaxDeepSeek34.8%
12Claude Opus 4.6Anthropic34.44%
13GPT-5 ProOpenAI31.64%
14GLM-5.131%
15DeepSeek-V3.2-SpecialeDeepSeek30.6%
16GLM-5Zhipu AI30.5%
17Kimi-K2.5Moonshot.AI30.1%
18Qwen3.5-397B-A17BAlibaba28.7%
19Step-3.5-Flash PaCoRe27.9%
20GPT-5.2OpenAI27.8%
21Gemma 4 31BGoogle26.5%
22GPT-5OpenAI25.32%
23GPT-5OpenAI25.3%
24Claude Opus 4.5Anthropic25.2%
25DeepSeek-V3.2DeepSeek25.1%
26GLM-4.7Zhipu AI24.8%
27Grok 4xAI24.5%
28Kimi K2.5Moonshot AI24.37%
29Qwen3.6-27B24%
30GPT-5.1OpenAI23.68%
31Step-3.5-Flash23.1%
32Gemini 2.5 ProGoogle21.64%
33Gemini 2.5 Pro21.6%
34Gemini 2.5 ProGoogle21.6%
35Qwen3.6-35B-A3B21.4%
36o3OpenAI20.32%
37GPT-5 miniOpenAI19.44%
38GPT-5 miniOpenAI19.4%
39MiniMax-M2.5MiniMaxAI19.4%
40Claude Opus 4.6Anthropic19%
41NVIDIA-Nemotron-3-Super-120B-A12B-BF1618.26%
42o4-miniOpenAI18.08%
43GLM-4.5Zhipu AI14.4%
44Claude Sonnet 4.5Anthropic13.72%
45Claude 4.5 SonnetAnthropic13.7%
46Claude Sonnet 4.6Anthropic13.2%
47Gemini 2.5 FlashGoogle12.1%
48Gemini 2.5 FlashGoogle12.08%
49Claude Opus 4.1Anthropic11.52%
50Gemini 2.5 Flash11%
51Claude Opus 4Anthropic10.72%
52GLM-4.5-AirZhipu AI10.6%
53NVIDIA-Nemotron-3-Nano-30B-A3B-BF1610.6%
54Gemini 3.1 Flash-LiteGoogle8.64%
55DeepSeek R1DeepSeek8.5%
56GLM-4.5Zhipu AI8.32%
57GLM-4.5-AirZhipu AI8.12%
58o1 ProOpenAI8.12%
59Claude 3.7 SonnetAnthropic8.04%
60o1OpenAI8%
61o1OpenAI7.96%
62Claude Sonnet 4Anthropic7.76%
63Gemini 2.0 Flash ThinkingGoogle6.56%
64Llama 4 MaverickMeta5.68%
65GPT-4.5 PreviewOpenAI5.44%
66GPT-4.1OpenAI5.4%
67Gemini 1.5 ProGoogle4.6%
68GPT-4.1 miniOpenAI4.6%
69Mistral-Medium-3Mistral4.52%
70Nova ProAmazon4.4%
71Claude 3.5 SonnetAnthropic4.08%
72Nova LiteAmazon3.64%
73GPT-4oOpenAI2.72%
74GPT-4oOpenAI2.7%
Source: agi.safe.ai · Leaderboard updated from live source.
§ 09 · Browse

By capability.

A shortcut into deeper leaderboards and per-task pages. All links resolve to live registry pages.

Capability
Reasoning
Multi-step, frontier-difficulty, GPQA and HLE.
Capability
Math
AIME 2025, olympiad-style short answer.
Capability
Code generation
LiveCodeBench, pass@1 on held-out contest problems.
Capability
Knowledge
MMLU and MMLU-Pro — breadth across 57 subjects.
Capability
Agentic
SWE-bench, Tau2-Bench, tool-use under real constraints.
Capability
All LLM datasets
The full index of text-in, text-out tasks.
§ 10 · Deep dives

By benchmark family.

Editorial pages with current rankings, eval methodology, and what the score actually means.

Deep dive
Benchmark ladder: 1M–7B
Which evals work at each model scale — from bits-per-byte and BLiMP to MMLU, GSM8K, and HumanEval.
Deep dive
Small English language models
A same-harness ranking plus a unified rerun roster for compact base models around the 125M class.
Deep dive
Coding benchmarks
LiveCodeBench, SWE-bench Verified, HumanEval+, MBPP — pass@1 across the coding leaderboards.
Deep dive
HumanEval & MBPP
The two saturating Python micro-benchmarks — what they still tell you and what they don’t.
Deep dive
Math benchmarks
AIME, MATH, Omni-MATH — frontier models on olympiad-style problems.
Deep dive
GSM8K
Grade-school math word problems — the canonical reasoning benchmark.
Deep dive
Reasoning benchmarks
GPQA, HLE, ARC-AGI — what frontier-difficulty actually means.
Deep dive
Open-weight models
Llama, Qwen, DeepSeek, Mistral — the open frontier vs the closed.
§ 11 · Related

Keep reading.

Adjacent sections of the registry.

Section
Agentic
SWE-bench, Terminal-Bench, tool-use and the trust problem in agent evals.
Section
Code generation
The pass@1 era and what comes after.
Section
Guide · Code models
Long-form guide comparing code models in production.
Section
News
Dated editorial notes when a benchmark moves.
Submit a result Read the methodology
Read next

Three places to go from here.

Sister hub
Code generation
SWE-bench, HumanEval, LiveCodeBench, Aider Polyglot — every code-generation benchmark and the harness behind it.
Sister hub
Agentic AI
Long-horizon agent benchmarks, OpenRouter adoption data, and which models actually show up in production agents.
Reference
Methodology
How scores are sourced, which sources count, and what we exclude. Required reading before quoting numbers.