I’m a Research Scientist at ServiceNow, where I work on large language models: reasoning for the Apriel models, multilingual instruction data with M2Lingual, and evaluation of audio LLMs with AU-Harness. Earlier, I contributed to the GEM benchmark for natural language generation.

A theme that runs through this work is measurement. Building models has become easier than knowing whether they are good: benchmarks saturate, and a high score doesn’t always mean a system will hold up in use. That gap — between what we can build and what we can verify — is one of the most interesting open problems in the field, and directly influences how and what to build next.

My PhD at UNC Charlotte focused on multi-party dialogue and online communication, where I saw how much data collection shapes what models can learn, and how much of language is behavior rather than text. These days, that background translates into an interest in closing the capability gap between English and other languages, extending evaluation to audio and visual reasoning, and understanding how AI systems behave once people actually use them.

My publications are on Google Scholar; you can also find me on LinkedIn and GitHub.