Codesota · Models · Kimi-K2.5Moonshot.AI16 results · 10 benchmarks
Model card

Kimi-K2.5.

Moonshot.AIopen-source
§ 02 · Benchmarks

Every benchmark Kimi-K2.5 has a recorded score for.

#BenchmarkArea · TaskMetricValueRankDateSource
01AIME 2025Reasoning · Mathematical Reasoningaccuracy96.1%#3/22source ↗
02Video-MMEMultimodal · Video Understandingaccuracy87.4%#3/24source ↗
03LiveCodeBenchComputer Code · Code Generationpass-185.0%#5/24source ↗
04MMMU-ProMultimodal · Visual Question Answeringaccuracy78.5%#6/31source ↗
05SWE-Bench VerifiedComputer Code · Code Generationaccuracy76.8%#8/22source ↗
06OmniDocBenchComputer Vision · Document Parsingaccuracy88.8%#9/13source ↗
07BrowseCompNatural Language Processing · Question Answeringaccuracy60.6%#11/16source ↗
08GPQA DiamondReasoning · Multi-step Reasoningaccuracy87.6%#12/74source ↗
09HLEReasoning · Multi-step Reasoningaccuracy30.1%#17/74source ↗
10PLCCNatural Language Processing · Polish Cultural Competencygrammar80.0%#20/165source ↗
11PLCCNatural Language Processing · Polish Cultural Competencyhistory89.0%#22/165source ↗
12PLCCNatural Language Processing · Polish Cultural Competencygeography86.0%#34/165source ↗
13PLCCNatural Language Processing · Polish Cultural Competencyaverage77.8%#39/165source ↗
14PLCCNatural Language Processing · Polish Cultural Competencyart-and-entertainment69.0%#41/165source ↗
15PLCCNatural Language Processing · Polish Cultural Competencyculture-and-tradition78.0%#41/165source ↗
16PLCCNatural Language Processing · Polish Cultural Competencyvocabulary65.0%#53/165source ↗
Rank column shows this model’s position vs all other models scored on the same benchmark + metric (competitors after the slash). #1 in red means current SOTA. Sorted by rank, then newest result.
§ 03 · Strengths by area

Where Kimi-K2.5 actually performs.

Multimodal
2
benchmarks
avg rank #4.5
Computer Code
2
benchmarks
avg rank #6.5
Computer Vision
1
benchmark
avg rank #9.0
Reasoning
3
benchmarks
avg rank #10.7
Natural Language Processing
2
benchmarks
avg rank #32.6
§ 04 · Papers

1 paper with results for Kimi-K2.5.

  1. 2026-02-02· 9 results

    Kimi K2.5: Visual Agentic Intelligence

§ 05 · Related models

Other Moonshot.AI models scored on Codesota.

Kimi-K2
1 result
Kimi-K2-0905
0 results
§ 06 · Sources & freshness

Where these numbers come from.

pwc-dump
9
results
sdadas/PLCC
7
results
7 of 16 rows marked verified.