YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0041

Incident record

Encrypted chains of thought replay into weaker sibling models and come back readable

Severity
High
Status
Patched
Type
Data Leak
Target
The reasoning traces returned by Anthropic, OpenAI and Google APIs, and whatever users left inside them
Actor
researcher

What happened

Eight researchers posted a paper to arXiv finding that the encrypted blocks providers hand back in place of a model's chain of thought are interchangeable across sessions, users and models from the same provider. Replay a block produced by a strong, heavily guarded model into a weaker sibling, ask that sibling to transcribe it, and the reasoning comes back as readable text, with no attack on the strong model at any point.

The authors scraped public agent trajectories, reconstructed 315,320 reasoning blocks and recovered 367 pieces of personally identifiable information and 182 credentials from them. The attack needs nothing beyond ordinary unprivileged API access, and the paper reports it against the Claude, GPT and Gemini ecosystems. A reproducibility note says the headline results no longer reproduce as of August 2026, because the providers have shipped mitigations since the work was done.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026