Codesota · Models · Gemini 3 ProGoogle13 results · 11 benchmarks
Model card

Gemini 3 Pro.

GoogleapiUndisclosed params1 current SOTA

Google flagship model.

§ 02 · Benchmarks

Every benchmark Gemini 3 Pro has a recorded score for.

#BenchmarkArea · TaskMetricValueRankDateSource
01GPQA DiamondReasoning · Multi-step Reasoningaccuracy91.9%#1/74source ↗
02LiveCodeBench ProComputer Code · Code Generationelo2439.00#2/10source ↗
03SWE-bench VerifiedAgentic AI · Autonomous Codingpct_resolved78.8%#2/3source ↗
04MMMU-ProMultimodal · Visual Question Answeringaccuracy80.0%#3/312026-01-15source ↗
05Tau2-BenchAgentic AI · Tool Usepass_rate69.0%#3/82025-11-18source ↗
06MMLUReasoning · Commonsense Reasoningaccuracy91.4%#6/642026-01-01source ↗
07HLEReasoning · Multi-step Reasoningaccuracy38.3%#6/74unverified
08SWE-benchComputer Code · Code Generationresolve-rate-agentic77.4%#8/252026-01-01source ↗
09OmniDocBenchComputer Vision · Document Parsingcomposite90.3%#8/34source ↗
10SWE-benchComputer Code · Code Generationresolve-rate77.4%#10/322026-01-01source ↗
11SWE-benchComputer Code · Code Generationresolve-rate-agentic76.2%#12/252025-12-01unverified
12SWE-Bench VerifiedComputer Code · Code Generationresolve-rate76.2%#12/39source ↗
13SWE-bench VerifiedAgentic AI · SWE-benchresolve-rate76.2%#19/81source ↗
Rank column shows this model’s position vs all other models scored on the same benchmark + metric (competitors after the slash). #1 in red means current SOTA. Sorted by rank, then newest result.
§ 03 · Strengths by area

Where Gemini 3 Pro actually performs.

Reasoning
3
benchmarks
avg rank #4.3 · 1 SOTA
Multimodal
1
benchmark
avg rank #3.0
Agentic AI
3
benchmarks
avg rank #8.0
Computer Vision
1
benchmark
avg rank #8.0
Computer Code
3
benchmarks
avg rank #8.8
§ 04 · Papers

1 paper with results for Gemini 3 Pro.

  1. 2023-10-10· Computer Code· 1 result

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.
§ 05 · Related models

Other Google models scored on Codesota.

Gemini 1.5 Pro
21 results
Gemini 2.5 Pro
16 results
Gemini-3.1-Pro
11 results
Gemma-3-27b
27B params · 11 results
Gemma 3 (27B, IT)
9 results
Gemma-2-27b-it
9 results
gemma-3-12b-it
9 results
gemma-3-1b-it
9 results
§ 06 · Sources & freshness

Where these numbers come from.

editorial
4
results
google-blog
2
results
livecodebench-pro-official
1
result
artificialanalysis.ai
1
result
codesota-shadow-mmlu
1
result
live-swe-agent
1
result
paddleocr-paper
1
result
swebench-leaderboard
1
result
google-internal
1
result
6 of 13 rows marked verified. · first result 2025-11-18, latest 2026-01-15.