MIT Technology Review named mechanistic interpretability one of its 10 Breakthrough Technologies for 2026, recognizing advances that map key features and pathways across AI models. In April 2026, Anthropic identified 171 emotion concept vectors in Claude Sonnet 4.5 that causally shift the model's behavior in the direction the emotion would predict. The clearest single advance of 2026 is documented in mechanistic interpretability breakthroughs, where probabilistic indicators meet actual measurements taken inside models. Rather than treating AI as a black box, researchers can now peer into the computational pathways that guide behavior, raising fresh questions about what internal states mean.
We built an emotion detector for machines and found emotions in the machine, but we still can't say whether the machine cares about having them.