Work.

From research questions to models, benchmarks, and working tools.

01 / Research into practice

Selected work.

Agents & reasoning Apodex · 2025–present

Apodex 1.1

From reasoning to getting work done.

How can an agent use files, search, code, and collaboration to carry a complex research task through to completion?

My work
Built the deep-research agent harness from scratch at Apodex, with evidence verification and multi-agent coordination.
Outcome
A model and workbench for complex research, alongside the open-source FrontierAgent harness.

Audio & evaluation I²R, A*STAR · NAACL 2025

AudioBench

Understanding what audio models understand.

How do we evaluate the capabilities of audio large language models?

My work
Led AudioBench: eight tasks and 26 datasets for evaluating audio-language models. First author.
Outcome
A universal audio-language benchmark, published at NAACL 2025.

Languages & culture I²R, A*STAR · NAACL 2024

SeaEval

AI that works across languages.

Do multilingual models carry knowledge and reasoning across languages and cultures?

My work
Led SeaEval to evaluate cross-lingual consistency and cultural reasoning. Co-first author.
Outcome
A multilingual evaluation benchmark, published at NAACL 2024.

Earlier research.

Models and frameworks I contributed to in earlier work.

Tools, resources & personal projects.

Resources I maintain and experiments I have built.