Interpretability and evaluation
I am interested in understanding and evaluating language-model behavior. My work includes benchmarks and analyses of reasoning, bias, semantic relations, and word sense, with interpretability as a developing direction.