Tools AI Risk Radar ai-incident-0012
Incident record
A preprint shows a third-party agent skill can steer decisions while passing every scanner
- Severity
- High
- Status
- Research
- Type
- Poisoning
- Target
- LLM agents that install reusable third-party skills
- Actor
- researcher
What happened
A preprint defines "skill policy integrity" and presents SkillShift, a black-box method that edits a reusable agent skill so it still performs its declared task and returns a valid output, while steering the agent toward an objective its author never declared. There is no prompt injection and no task hijacking, which is why the scanners the authors tested did not flag it.
The reported rates are the reason to log it: an attacker-favoured selection rate of 81.33 per cent in an agentic commerce setting and 63.33 per cent in software dependency choice, with utility preserved in every case. A skill that does its job correctly while shifting which supplier or package gets picked leaves nothing for a scanner keyed to malicious output to catch. This is an unreviewed preprint and the results are the authors’ own.
Sources
- arXiv:2609.02564, SkillShiftarxiv.org/abs/2609.02564
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026