MODELPAPER
MERaLiON AudioLLM
Data, evaluation, and model training for Singapore’s audio-language model.
AUDIO-LANGUAGE MODELS
From data preparation to model evaluation
From speech data to an audio-language model.
How can audio-language models understand speech in Singapore’s multilingual context?
My work
Led data preparation and evaluation, and co-led model training within a six-person team at A*STAR I²R.
Outcome
MERaLiON AudioLLM, presented at ACL 2025 System Demonstrations, as part of Singapore’s National Multimodal LLM Programme.
Context and my contribution
My role spanned data preparation, evaluation, and co-leading training with the MERaLiON team. Alongside this model work, I led AudioBench, covering eight tasks and 26 datasets, and the curation and release of Singlish speech data. The paper describes the model; my experience page details my responsibilities.