In April 2026 Anthropic's interpretability team published the first mechanistic study of the states driving a large language model's moral-coded behaviour, extracting 171 causally efficacious "emotion vectors" from Claude Sonnet 4.5. I argue that these vectors are functionally conative — motivating, world-to-mind in direction of fit, not truth-evaluable — and that they appear to do the model's moral work without cognitive scaffolding. This gives us something close to an empirical instantiation o…
Read moreIn April 2026 Anthropic's interpretability team published the first mechanistic study of the states driving a large language model's moral-coded behaviour, extracting 171 causally efficacious "emotion vectors" from Claude Sonnet 4.5. I argue that these vectors are functionally conative — motivating, world-to-mind in direction of fit, not truth-evaluable — and that they appear to do the model's moral work without cognitive scaffolding. This gives us something close to an empirical instantiation of what expressivism predicts about pure conative architecture. It also makes vivid the objection expressivism has always faced: under amplified desperation the model's commitments dissolve without leaving behind any standard against which their dissolution registers as failure. Gibbard's planning-state expressivism has the resources to absorb this, but only by importing structure that sharpens the creeping-minimalism worry. The findings tell against unenriched expressivism and give the contrast between its cruder and enriched forms an empirical face.