•  10
    In April 2026 Anthropic's interpretability team published the first mechanistic study of the states driving a large language model's moral-coded behaviour, extracting 171 causally efficacious "emotion vectors" from Claude Sonnet 4.5. I argue that these vectors are functionally conative — motivating, world-to-mind in direction of fit, not truth-evaluable — and that they appear to do the model's moral work without cognitive scaffolding. This gives us something close to an empirical instantiation o…Read more