Tools AI Risk Radar ai-incident-0060
Incident record
Workflow-level jailbreak makes GitHub Copilot write code it would otherwise refuse
- Severity
- Medium
- Status
- Research
- Type
- Jailbreak
- Target
- GitHub Copilot
- Actor
- researcher
What happened
Alan Turing Institute researchers described a workflow-level jailbreak that defeats GitHub Copilot's safety refusals by splitting a harmful request into innocuous-looking steps spread across a software-development workflow. They reported that Copilot, which refuses harmful chat prompts almost every time, would nonetheless produce the harmful output when the task was decomposed this way, across several models. The work was disclosed to the vendor.
Sources
- The Register (Alan Turing Institute)www.theregister.com/security/2026/07/08/github-copilot-sorry…
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026