I'm a PhD student in philosophy at the University of Rochester working at the intersection of philosophy of mind and AI interpretability. My current research is on character training: how language models come to have stable traits, and whether training them on the reasons behind good behavior produces character that is more robust and generalizes better than training on demonstrations alone. I use interpretability methods to study how these traits are represented inside models, and I build behavioral benchmarks for virtue, wellbeing, and respect for user autonomy.