YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0060

Incident record

Workflow-level jailbreak makes GitHub Copilot write code it would otherwise refuse

Severity
Medium
Status
Research
Type
Jailbreak
Target
GitHub Copilot
Actor
researcher

What happened

Alan Turing Institute researchers described a workflow-level jailbreak that defeats GitHub Copilot's safety refusals by splitting a harmful request into innocuous-looking steps spread across a software-development workflow. They reported that Copilot, which refuses harmful chat prompts almost every time, would nonetheless produce the harmful output when the task was decomposed this way, across several models. The work was disclosed to the vendor.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026