Researchers Steal Hidden Reasoning from Anthropic, OpenAI, and Google LLMs via Replay Attack
Stealing Reasoning Traces from Proprietary LLM APIs
A new paper reveals that proprietary LLM APIs from Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks that can be replayed across sessions, users, and models. By injecting a trace from a frontier model into a weaker, jailbroken sibling, the researchers recover the stronger model's hidden reasoning in plaintext, without attacking the stronger model directly or triggering anti-distillation safeguards. The attack works on claude-opus-4-8, GPT-4, and Gemini models.
Proprietary reasoning can be recovered from its encrypted traces.