Codesota · Benchmark · VoiceBenchHome/Leaderboards/Multimodal Media/Audio + Text to Text/VoiceBench
NUS, Alibaba, Tsinghua

VoiceBench.

VoiceBench is a multi-facet evaluation suite for LLM-based voice assistants, covering general knowledge, instruction following, safety refusal, and robustness to speaker accents and background noise across diverse speech inputs.

Paper Leaderboard
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result Report issue

Overall Score

Overall Score is the reported evaluation metric for VoiceBench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Overall Scoreverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Ultravox-GLM-4P7
VoiceBench overall-score. Rank #1 on VoiceBench leaderboard as of 2026-03-28. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified88.862026Source ↗Looks wrong?
02Whisper-v3-large + GPT-4o (cascade)
VoiceBench overall-score. Cascade baseline: Whisper-large-v3 + GPT-4o. Rank #3. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified87.82026Source ↗Looks wrong?
03GPT-4o-Audio
VoiceBench overall-score. Rank #5 on VoiceBench leaderboard as of 2026-03-28. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified86.752026Source ↗Looks wrong?
04Whisper-v3-large + LLaMA-3.1-8B (cascade)
VoiceBench overall-score. Cascade baseline from original paper. Rank #9. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified77.482026Source ↗Looks wrong?
05Kimi-Audio
VoiceBench overall-score. Rank #10 on VoiceBench leaderboard as of 2026-03-28. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified76.912026Source ↗Looks wrong?
06MiniCPM-o
VoiceBench overall-score. Rank #15 on VoiceBench leaderboard as of 2026-03-28. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified71.232026Source ↗Looks wrong?
07VITA-1.5
VoiceBench overall-score. Rank #19 on VoiceBench leaderboard as of 2026-03-28. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified64.532026Source ↗Looks wrong?
08Qwen2-Audio
VoiceBench overall-score. Rank #27. From original paper Table 3 and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified55.82026Source ↗Looks wrong?
09LLaMA-Omni
VoiceBench overall-score. Rank #34. From original paper Table 3 and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified41.122026Source ↗Looks wrong?
10VITA-1.0
VoiceBench overall-score. Rank #35. From original paper Table 3 (as VITA) and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified36.432026Source ↗Looks wrong?
11Mini-Omni2
VoiceBench overall-score. Rank #37. From original paper Table 3 and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified33.492026Source ↗Looks wrong?
12Mini-Omni
VoiceBench overall-score. Rank #38. From original paper Table 3 and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified30.422026Source ↗Looks wrong?
13Moshi
VoiceBench overall-score. Rank #39 (last). From original paper Table 3 and leaderboard. Source: matthewcym.github.io/VoiceBench/ (accessed 2026-03-28) and arXiv:2410.17196v2.
verified29.512026Source ↗Looks wrong?
§ 04 · Submit a result

Add to the leaderboard.

← Back to Audio + Text to Text