MIT Technology Review named mechanistic interpretability one of its 10 Breakthrough Technologies for 2026, marking a turning point in how scientists approach the question of what happens inside AI models. Mechanistic interpretability is the study of how large language models work internally, reverse-engineering the features and computational pathways a model uses to turn a prompt into a response. The breakthrough has immediate stakes for consciousness research: detecting consciousness-like internal states first requires being able to read what a model is actually doing inside the black box. Anthropic's emotion vectors paper identified 171 emotion concept vectors in Claude Sonnet 4.5 that causally shift the model's behavior in the direction the emotion would predict, showing the field can now produce "concretely actionable findings." OpenAI is building what they term an "AI lie detector" using model internals to identify when models are being deceptive, examining internal representations to determine whether the model's internal state corresponds to the truth or contradicts it.
Knowing how a machine thinks turns out to be harder than knowing that it does, and now that researchers are learning to look inside, they're finding it's possible a system can seem conscious while containing nothing resembling experience.