BENCHMARKPAPER
AudioBench
A universal benchmark for audio large language models. NAACL 2025.
AUDIO MODEL EVALUATION
From input to measurable behavior
Understanding what audio models understand.
How do we evaluate the capabilities of audio large language models?
My work
Led AudioBench: eight tasks and 26 datasets for evaluating audio-language models. First author.
Outcome
A universal audio-language benchmark, published at NAACL 2025.
Context and my contribution
Audio-language models need to be assessed across different tasks, rather than through a single demonstration. I led AudioBench to bring eight tasks and 26 datasets into a common benchmark. The paper and code make the evaluation setup available to other researchers. This complemented my data preparation and evaluation work on MERaLiON AudioLLM.