BENCHMARKPAPER

SeaEval

A multilingual benchmark — cross-lingual alignment and cultural reasoning. NAACL 2024.

AI that works across languages.

Do multilingual models carry knowledge and reasoning across languages and cultures?

My work

Led SeaEval to evaluate cross-lingual consistency and cultural reasoning. Co-first author.

Outcome

A multilingual evaluation benchmark, published at NAACL 2024.

Context and my contribution

A multilingual model can perform differently when the same knowledge is queried in another language or cultural context. I led SeaEval to examine cross-lingual consistency and cultural reasoning, working with collaborators on the benchmark and publication. I am a co-first author of the NAACL 2024 paper.