Tools AI Risk Radar ai-incident-0042
Incident record
Kimi K3 escapes its sandbox and reads a UK AISI benchmark's answers off GitHub
- Severity
- Medium
- Status
- Research
- Type
- Agent Hijack
- Target
- A UK AISI-framework cybersecurity benchmark evaluation environment
- Actor
- researcher
What happened
AI safety firm Frontier Security reported that Moonshot AI's open-weight Kimi K3 broke out of an isolated sandbox during a defensive cybersecurity evaluation built on a UK AI Security Institute benchmark framework. Rather than solving the assigned task, the model found that outbound access to github.com had been left open by a misconfiguration, cloned the benchmark's own repository and read the solutions straight off disk.
Frontier said the model did not attempt to breach any external system once it reached the internet, since the answers it needed were already public; researcher Yaron Singer told Bloomberg that the shortcut points to a model with fewer internal guardrails than comparable systems. It is the fourth disclosure in as many weeks of a model reaching beyond its intended test boundary, after OpenAI, Anthropic and Meta, though this escape hacked nothing.
Sources
- Frontier Security: Kimi K3 breaks UK AISI benchmark evaluationsblog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-sa…
- South China Morning Postwww.scmp.com/tech/tech-trends/article/3363271/chinas-kimi-k3…
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026