MSc in Brain & Cognitive Sciences, 2024
University of Amsterdam
Bsc in Computational Neuroscience, 2022
University of Southern California
My research at the intersection of computational cognitive science and explainable AI investigates the alignment of causal reasoning in humans and language models. My work explores whether language models faithfully replicate human judgments in causal reasoning tasks. In addition to these behavioral evaluations of LLMs, my research examines the internal mechanisms underlying causal reasoning in language models by implementing methods from mechanistic interpretability, such as activation patching, SAEs and activation steering. I am also interested in modeling the communication of causal information via multi-agent RL.