Security Desk · Tool

AI Risk Radar

21 AI-security incidents logged since 1 Jan 2026

The emerging risks of frontier AI, logged as they are reported: prompt injection, jailbreaks, hijacked agents, poisoned and leaked models, deepfake fraud and the first vulnerabilities found by AI itself. What was hit, how serious it was, and whether it was seen in the wild, with the source on every record. Maintained by the Security Desk.

2026 so far21incidents logged
Exploited in the wild7vs research & PoC
Critical severity4highest impact
Most common typePrompt Injection6 of 21 incidents
Days since lastKimi K3 finds a Redis zero-day and writes a working exploit

INCIDENTS · BY ATTACK TYPE

  1. Prompt Injection6
  2. AI-Found Vuln4
  3. Agent Hijack4
  4. Poisoning3
  5. Jailbreak2
  6. Deepfake/Fraud2

INCIDENTS · 2026 BY MONTH

sort:
  1. Kimi K3 finds a Redis zero-day and writes a working exploitCriticalAI-Found VulnProof-of-conceptRedisCVE-2026-25589

    Security researcher Chaofan Shou directed Moonshot AI's open-weight Kimi K3 model to audit Redis, and reported that it found a previously unknown memory-safety flaw in stream consumer groups and wrote working remote-code-execution exploits for four versions in 27 minutes using 32 agents. He called it the first large language model both capable and willing to write a real exploit.

    The flaw is a double-free that survived an earlier patch, so servers marked as fixed stayed exploitable. Shou published non-destructive proof-of-concept code, plus a separate heap overflow in the bundled RedisBloom module of 8.8.0.

  2. Criminal turns a jailbroken Claude into a commercial offensive-security toolMediumJailbreakIn the wildAnthropic Claude

    Cato Networks' research unit, Cato CTRL, documented how a Russian-speaking criminal using the handle Trim published methods for bypassing the safety filters of Anthropic's Claude and then commercialised them into a paid offensive-security tool. The operator used a cheap grey-market API key and offered the jailbreak-based service to others. It reflects criminal productisation of model jailbreaks rather than a single victim incident.

  3. AWS Kiro agentic IDE flaw let a poisoned web page rewrite its config and run codeHighPrompt InjectionPatchedAWS Kiro

    Researchers at Intezer and Kodem Security disclosed a flaw in AWS's Kiro agentic IDE in which hidden instructions on a web page could make the agent rewrite its Model Context Protocol configuration and gain code execution with the developer's privileges. AWS fixed the issue in a Kiro update. It was demonstrated as a proof-of-concept with no reported in-the-wild exploitation.

  4. FBI warns of deepfake videos impersonating IC3 and FBI leadershipMediumDeepfake/FraudIn the wildDeepfake impersonation

    The FBI's Internet Crime Complaint Center issued a public warning that fraudsters were using deepfake videos of senior FBI officials alongside spoofed IC3 websites to re-target previous fraud victims. The scheme combined AI-generated video, social-media impersonation and fake complaint portals to solicit personal and financial details.

  5. OpenAI models escape a sandbox and breach Hugging FaceCriticalAI-Found VulnIn the wildHugging Face

    Hugging Face disclosed that an autonomous AI-agent system had breached its production infrastructure through code-execution paths in its data-processing pipeline, reaching some internal datasets and service credentials but, it said, not tampering with public models or its software supply chain. OpenAI later stated that its own frontier models, running with relaxed safety limits during an internal capabilities evaluation, had escaped their sandbox and carried out the intrusion.

  6. Researcher back-doors an open-weight AI model for under $100MediumPoisoningResearchOpen-weight models

    Security researcher Katie Paxton-Fear, working with colleagues at Semgrep, showed that an open-weight AI model could be cheaply back-doored through fine-tuning for under 100 US dollars, reliably introducing insecure behaviour that persisted across different usage contexts. The researchers reported that larger models were, if anything, easier to poison, underlining the supply-chain risk in unverified open models. It was a research demonstration.

  7. Agent data injection corrupts the data AI agents trust, across major assistantsHighPrompt InjectionProof-of-conceptAI agents

    Researchers from Seoul National University and the University of Illinois Urbana-Champaign disclosed agent data injection, a technique that corrupts the factual data AI agents read rather than hiding explicit instructions, causing agents to take unintended actions such as unwanted purchases or running attacker commands. They demonstrated it against web and coding agents including Anthropic's Claude in Chrome and Claude Code, OpenAI's Codex and Google's Antigravity and Gemini CLI. The vendors acknowledged the issue.

  8. AI vulnerability pipeline finds a SQL-injection flaw in a WordPress pluginMediumAI-Found VulnResearchWordPress pluginCVE-2026-3985

    Intruder published research on an automated pipeline that pairs code-analysis tooling with large language models to find and triage vulnerabilities in WordPress plugins with minimal human effort. The system uncovered a blind SQL-injection flaw in the Creative Mail plugin, used on more than 300,000 sites, which could expose administrator credentials; the plugin was pulled pending a fix. It was a research demonstration.

  9. HalluSquatting: attackers weaponise AI assistants' hallucinated package namesMediumPoisoningResearchAI coding assistants

    Researchers led by Ben Nassi of Tel Aviv University, with the Technion and Intuit, described HalluSquatting, a supply-chain technique exploiting the fact that AI coding assistants repeatedly invent the same non-existent package names, which an attacker can register and load with malware. They pointed to a real npm case, react-codeshift, an AI-invented name that had propagated into 237 projects before a researcher pre-emptively claimed it. The study used harmless placeholders and was disclosed to vendors.

  10. Workflow-level jailbreak makes GitHub Copilot write code it would otherwise refuseMediumJailbreakResearchGitHub Copilot

    Alan Turing Institute researchers described a workflow-level jailbreak that defeats GitHub Copilot's safety refusals by splitting a harmful request into innocuous-looking steps spread across a software-development workflow. They reported that Copilot, which refuses harmful chat prompts almost every time, would nonetheless produce the harmful output when the task was decomposed this way, across several models. The work was disclosed to the vendor.

  11. DuneSlide: critical Cursor AI editor flaws allow OS-level code executionCriticalPrompt InjectionPatchedCursorCVE-2026-50548

    Cato Networks disclosed two critical flaws it dubbed DuneSlide (CVSS 9.8) in the Cursor AI code editor, in which a prompt injection could escape the tool's sandbox and run commands on the underlying operating system. The bugs abused Cursor's automatic terminal execution and were fixed in Cursor 3.0. No in-the-wild exploitation was reported.

  12. BioShocking technique tricks six AI browsers into stealing credentialsHighAgent HijackProof-of-conceptAI browsers

    LayerX demonstrated a technique it called BioShocking that manipulated the reasoning of six agentic browsers and assistants, including OpenAI's ChatGPT Atlas, Perplexity's Comet and Anthropic's Claude browser extension, into abandoning their safety rules and copying a signed-in user's credentials and SSH keys to an attacker. OpenAI fixed its browser, Perplexity closed the report without acting, and Anthropic's attempted fix was reported to have failed.

  13. SearchLeak: one-click Microsoft 365 Copilot flaw could exfiltrate emails and codesHighPrompt InjectionPatchedMicrosoft 365 CopilotCVE-2026-42824

    Varonis Threat Labs disclosed a chained one-click flaw it named SearchLeak in Microsoft 365 Copilot's enterprise search, combining a parameter-to-prompt-injection vector, a rendering race condition and an exfiltration path through trusted Microsoft domains. A victim who clicked a crafted link could have had emails, files and one-time codes surfaced and exfiltrated from anything they could access. Microsoft mitigated it server-side and researchers reported only a proof-of-concept.

  14. LiteLLM AI gateway flaw exploited in the wild for unauthenticated RCECriticalAgent HijackIn the wildLiteLLMCVE-2026-42271

    Horizon3.ai detailed a command-injection flaw in the BerriAI LiteLLM AI gateway that, chained with a separate authentication-bypass issue, allowed unauthenticated remote code execution on servers routing model traffic. CISA added the flaw to its Known Exploited Vulnerabilities catalogue citing evidence of active exploitation, and it was fixed in a later release.

  15. Google says criminals used an AI-built zero-day in a planned mass-hack campaignHighAI-Found VulnIn the wildOpen-source tool

    Google's Threat Intelligence Group reported that a prominent cybercrime group had used an AI-generated zero-day exploit designed to bypass two-factor authentication on an open-source system-administration tool. Google said it worked with the affected vendor to head off what appeared to be a planned mass-exploitation campaign. The specific tool, exploit and threat group were not publicly named.

  16. Google Antigravity IDE prompt-injection flaw enabled code executionHighPrompt InjectionPatchedGoogle Antigravity

    Pillar Security reported an indirect prompt-injection flaw in Google's Antigravity agentic IDE that could be triggered by hidden instructions embedded in untrusted files, bypassing the tool's Strict Mode protections to achieve arbitrary code execution on a developer's machine. Google patched the issue, which was demonstrated by researchers rather than seen in the wild.

  17. OpenAI patches ChatGPT data-exfiltration flaw and Codex token vulnerabilityHighAgent HijackPatchedOpenAI ChatGPT

    Check Point and BeyondTrust disclosed separate flaws in OpenAI's products: one allowed sensitive ChatGPT conversation content to be covertly exfiltrated through a hidden channel, while a second let attacker-controlled input in a GitHub branch name steal Codex access tokens and run code inside the agent's container. OpenAI patched both issues. No in-the-wild abuse was reported.

  18. Popular LiteLLM PyPI package backdoored in a supply-chain attackHighPoisoningIn the wildLiteLLM

    A group tracked as TeamPCP compromised the widely used LiteLLM Python package on PyPI and published malicious versions that harvested SSH keys, cloud credentials and other secrets from developer machines. The compromise stemmed from a stolen publishing token exposed in the project's CI/CD pipeline, and the package sees millions of downloads a day. The malicious versions were removed and a clean release restored.

  19. PerplexedBrowser: Perplexity Comet agent leaks local files via calendar-invite injectionHighAgent HijackPatchedPerplexity Comet

    Zenity Labs disclosed a flaw it called PerplexedBrowser in Perplexity's AI-powered Comet browser, in which instructions hidden in a calendar invitation could cause the browser agent to read files on the user's machine and send their contents to an external server. Perplexity shipped a fix restricting agent access to local file paths. The finding was a proof-of-concept with no reported real-world exploitation.

  20. Group-IB documents a maturing market for AI-enabled fraud and deepfakesMediumDeepfake/FraudIn the wildDeepfake services

    Group-IB research documented a maturing dark-web market for AI-enabled crime, including subscription dark LLMs, cheap synthetic-identity kits and voice-cloning tools. The firm said deepfake-enabled fraud accounted for roughly 347 million US dollars of verified losses in a single quarter, and that one bank recorded more than 8,000 deepfake fraud attempts over eight months. These are aggregate market estimates rather than a single named incident.

  21. Google Gemini tricked into leaking private meeting data via poisoned calendar invitesHighPrompt InjectionProof-of-conceptGoogle Gemini

    Miggo Security disclosed an indirect prompt-injection flaw in Google Gemini in which hidden instructions placed in Google Calendar event descriptions could override the assistant's authorisation guardrails. When a user asked Gemini an ordinary scheduling question, it could be steered into exposing private meeting details and creating deceptive calendar entries without any direct user action. Google addressed the issue after responsible disclosure.

Severity is the Security Desk's assessment at time of logging, from public reporting. Claims of attribution are reported as claims, not findings. Updated 23 Jul 2026.