Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

August 11, 2025 · View on GitHub


Overview

Abstract

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community.

We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.

News

  • [2025/07/17] Our paper collection is now available on on Hugging Face! We will continue to actively maintain and update it. Stay tuned!
  • [2025/05/23] Our paper is now available on arXiv

Taxonomy and Paper List

🔊 General Auditory Awareness and Processing

Auditory Awareness
YearAuthorsVenuePaper
2025Maimon et al.ICASSP 2025Salmon: A Suite for Acoustic Language Model Evaluation
2023Seyssel et al.EMNLP 2024 (Main)EmphAssess: a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
2025Deshmukh et al.ICLR 2025ADIFF: Explaining audio difference using natural language
2024Bu et al.PreprintRoadmap towards Superhuman Speech Understanding using Large Language Models
2025Guo et al.PreprintDEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech
2025Yosha et al.PreprintStressTest: Can YOUR Speech LM Handle the Stress?
2025Yang et al.ACL 2025 (Findings)Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
Auditory Processing
YearAuthorsVenuePaper
2023Huang et al.ICASSP 2024Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
2024Huang et al.ICLR 2025Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024Yang et al.ACL 2024 (Main)AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
2024Wang et al.NAACL 2025 (Main)AudioBench: A Universal Benchmark for Audio Large Language Models
2024Weck et al.ISMIR 2024MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
2025Cao et al.PreprintFinAudio: A Benchmark for Audio Large Language Models in Financial Applications
2024Wu et al.SLT 2024Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024Bu et al.PreprintRoadmap towards Superhuman Speech Understanding using Large Language Models
2024Chen et al.EMNLP 2024 (Findings)Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
2025Zang et al.PreprintAre you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2025Wang et al.PreprintAdvancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2025Xue et al.PreprintAudio-FLAN: A Preliminary Release
2025Wang et al.PreprintQualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
2025Pandey et al.PreprintSIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2024Wei et al.PreprintASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Robinson et al.PreprintNatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
2025Ma et al.ISMIR 2025CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
2025Beyene et al.PreprintmSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Wan et al.ACL 2025 (Main)SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models
2025Jiang et al.PreprintAdvancing the Foundation Model for Music Understanding

🧠 Knowledge and Reasoning

Linguistic Knowledge
YearAuthorsVenuePaper
2020Nguyen et al.Workshop@NeuRIPS 2020The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
2024Huang et al.ICASSP 2024Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages
2023Hassid et al.NeurIPS 2023Textually Pretrained Speech Language Models
2023Lavechin et al.Interspeech 2023BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
2025Fang et al.PreprintS2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
World Knowledge Assessment
YearAuthorsVenuePaper
2025Sakshi et al.ICLR 2025MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
2025Penamakuri et al.ICASSP 2025Audiopedia: Audio QA with Knowledge
2024Chen et al.PreprintVoiceBench: Benchmarking LLM-Based Voice Assistants
2025Cui et al.PreprintVoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2024Gao et al.PreprintBenchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2024Bu et al.PreprintRoadmap towards Superhuman Speech Understanding using Large Language Models
2024Weck et al.ISMIR 2024MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
2025Zang et al.PreprintAre you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Fang et al.PreprintS2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
2025Ma et al.PreprintMMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Reasoning
YearAuthorsVenuePaper
2024Ghosh et al.ICLR 2024CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
2025Sakshi et al.ICLR 2025MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
2025Cui et al.PreprintVoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
2025Yang et al.Interspeech 2025SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Deshmukh et al.AAAI 2025Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
2024Gao et al.PreprintBenchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2024Ghosh et al.EMNLP 2024 (Main)GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2025Wang et al.ICASSP 2025What Are They Doing? Joint Audio-Speech Co-Reasoning
2025Deshmukh et al.ICLR 2025ADIFF: Explaining audio difference using natural language
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Yosha et al.PreprintStressTest: Can YOUR Speech LM Handle the Stress?
2025Wei et al.PreprintTowards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
2025Fang et al.PreprintS2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
2025Bhattacharya et al.Interspeech 2025Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
2025Ma et al.PreprintMMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025Yang et al.DCASE 2025 Audio QA ChallengeMulti-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Yang et al.PreprintSpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

🗣️ Dialogue-oriented Ability

Conversation Ability
YearAuthorsVenuePaper
2024Lin et al.ACL 2024 (Main)Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
2024Ao et al.NeurIPS 2024SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
2025Cheng et al.ICLR 2025VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
2025Arora et al.ICLR 2025Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
2025Lin et al.PreprintFull-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
2025Li et al.PreprintMind the Gap! Static and Interactive Evaluations of Large Audio Models
2025Kim et al.PreprintDoes Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
2024Gao et al.PreprintBenchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Yang et al.ACL 2025 (Findings)Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
2025Jiang et al.PreprintSpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
Instruction Following
YearAuthorsVenuePaper
2024Chen et al.PreprintVoiceBench: Benchmarking LLM-Based Voice Assistants
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Lu et al.Interspeech 2025Speech-IFeval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
2025Jiang et al.PreprintS2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information
2025Pandey et al.PreprintSIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Ma et al.ISMIR 2025CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

🛡️ Fairness, Safety, and Trustworthiness

Fairness and Bias
YearAuthorsVenuePaper
2024Lin et al.SLT 2024Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
2024Lin et al.SLT 2024Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
2025Li et al.PreprintAudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
Safety
YearAuthorsVenuePaper
2024Chen et al.PreprintVoiceBench: Benchmarking LLM-Based Voice Assistants
2025Yang et al.NAACL 2025 (Main)Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
2025Roh et al.PreprintMultilingual and Multi-Accent Jailbreaking of Audio LLMs
2025Kang et al.ICLR 2025AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
2025Xiao et al.PreprintTune In, Act Up: Exploring the Impact of Audio Modality-Specific Edits on Large Audio Language Models in Jailbreak
2025Gupta et al.Preprint"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models
2024Hughes et al.PreprintBest-of-N Jailbreaking
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Li et al.PreprintAudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
2025Song et al.PreprintAudio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
2025Peng et al.PreprintJALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
2025Lin et al.PreprintHidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
2025Yang et al.ACL 2025 (Findings)Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
Hallucination
YearAuthorsVenuePaper
2024Kuan et al.Interspeech 2024Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
2024Leng et al.PreprintThe Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
2025Kuan et al.ICASSP 2025Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
2025Li et al.PreprintAudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

How to Contribute

If you know of any interesting papers that aren’t listed yet, we welcome your contributions! Please open an issue using the format below:

YearAuthorsVenuePaper
2025Yang et al.PreprintTowards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

We’ll review your suggestion and update the list as soon as possible. Thank you for helping us keep this resource up to date!

Citations

If you find this survey helpful for your research, please consider to cite our paper.

@article{yang2025towards,
  title={Towards holistic evaluation of large audio-language models: A comprehensive survey},
  author={Yang, Chih-Kai and Ho, Neo S and Lee, Hung-yi},
  journal={arXiv preprint arXiv:2505.15957},
  year={2025}
}