LLMs Can Introspect on Their Own Internal States, Study Finds
Emergent Introspective Awareness in Large Language Models

A new study from Anthropic probes whether large language models can introspect on their internal states. By injecting known concepts into model activations and measuring self-reported states, researchers found that models can sometimes notice and identify these injections, recall prior representations, and even distinguish their own outputs from artificial prefills. Claude Opus 4 and 4.1 showed the greatest introspective awareness, though reliability varies. The results suggest current LLMs possess limited but real introspective capabilities.
We stress that in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.