Mechanistic interpretability aims to map key features and pathways across AI models. In 2024, Anthropic announced it had built a microscope that let researchers peer inside Claude and identify features corresponding to recognizable concepts like Michael Jordan and the Golden Gate Bridge. The field accelerated in 2025 and 2026. In April 2026, Anthropic's emotion vectors paper identified 171 emotion concept vectors in Claude Sonnet 4.5 that causally shift the model's behavior in the direction the emotion would predict. MIT Technology Review recognized mechanistic interpretability as one of the 10 Breakthrough Technologies for 2026.
It turns out you can ask a machine what it's actually thinking, provided you build the right tools to listen.