Avg Score is the reported evaluation metric for AudioBench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Trust tiers for Avg Scoreverifiedpapervendorcommunityunverified
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
Average of 19 non-ASR higher-is-better subtasks from AudioBench Table 2 (arXiv:2406.16020v5). WavLLM. Note: paper evaluates 5 models; GPT-4o/Gemini not included in original evaluation.
Average of 19 non-ASR higher-is-better subtasks from AudioBench Table 2 (arXiv:2406.16020v5). SALMONN. Note: paper evaluates 5 models; GPT-4o/Gemini not included in original evaluation.
Average of 19 non-ASR higher-is-better subtasks from AudioBench Table 2 (arXiv:2406.16020v5). Qwen2-Audio-Instruct. Note: paper evaluates 5 models; GPT-4o/Gemini not included in original evaluation.
Average of 19 non-ASR higher-is-better subtasks from AudioBench Table 2 (arXiv:2406.16020v5). Whisper+LLaMA-3 (cascade). Note: paper evaluates 5 models; GPT-4o/Gemini not included in original evaluation.
Average of 19 non-ASR higher-is-better subtasks from AudioBench Table 2 (arXiv:2406.16020v5). Qwen-Audio-Chat. Note: paper evaluates 5 models; GPT-4o/Gemini not included in original evaluation.