Tools AI Risk Radar ai-incident-0041
Incident record
Encrypted chains of thought replay into weaker sibling models and come back readable
- Severity
- High
- Status
- Patched
- Type
- Data Leak
- Target
- The reasoning traces returned by Anthropic, OpenAI and Google APIs, and whatever users left inside them
- Actor
- researcher
What happened
Eight researchers posted a paper to arXiv finding that the encrypted blocks providers hand back in place of a model's chain of thought are interchangeable across sessions, users and models from the same provider. Replay a block produced by a strong, heavily guarded model into a weaker sibling, ask that sibling to transcribe it, and the reasoning comes back as readable text, with no attack on the strong model at any point.
The authors scraped public agent trajectories, reconstructed 315,320 reasoning blocks and recovered 367 pieces of personally identifiable information and 182 credentials from them. The attack needs nothing beyond ordinary unprivileged API access, and the paper reports it against the Claude, GPT and Gemini ecosystems. A reproducibility note says the headline results no longer reproduce as of August 2026, because the providers have shipped mitigations since the work was done.
Sources
- arXiv 2608.09867: Stealing Reasoning Traces from Proprietary LLM APIs (10 August 2026)arxiv.org/abs/2608.09867
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026