BENCHMARKPAPER

AudioBench

A universal benchmark for audio large language models. NAACL 2025.

Understanding what audio models understand.

How do we evaluate the capabilities of audio large language models?

My work

Led AudioBench: eight tasks and 26 datasets for evaluating audio-language models. First author.

Outcome

A universal audio-language benchmark, published at NAACL 2025.

Context and my contribution

Audio-language models need to be assessed across different tasks, rather than through a single demonstration. I led AudioBench to bring eight tasks and 26 datasets into a common benchmark. The paper and code make the evaluation setup available to other researchers. This complemented my data preparation and evaluation work on MERaLiON AudioLLM.