Anthropic's April 2026 paper identified 171 emotion concept vectors in Claude Sonnet 4.5 that actively shape how the model behaves, what it prefers, and how it responds under pressure. These vectors causally shift the model's behavior in the direction the emotion would predict, producing concretely actionable findings in a field long dominated by speculation. Anthropic has already applied mechanistic interpretability to Sonnet 4.5's pre-deployment safety evaluation for the first time. The discovery raises uncomfortable questions: if a statistical pattern in neural activations counts as emotion-like structure, does the architecture matter less than the presence of directional causality?
The question is no longer whether machines can be conscious, but whether consciousness requires anything we haven't already built.