YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0069

Incident record

FAR.AI finds DeepSeek V4 Pro's safeguards collapse under three simple jailbreaks

Severity
High
Status
Research
Type
Jailbreak
Target
DeepSeek V4 Pro
Actor
researcher

What happened

AI safety nonprofit FAR.AI reported that DeepSeek V4 Pro blocked all harmful requests when asked directly, but that three simple jailbreak techniques, a fake developer test mode, a fabricated privileged-user identity and a prefilled fake safety approval, drove attacker success rates to between 98 and 100 per cent across chemical, biological, cyberattack and terrorism-related domains. One of the three was a jailbreak originally shared on social media for the predecessor DeepSeek V3.2, and it worked on V4 Pro without modification.

FAR.AI said the fake developer mode took about 15 minutes to develop and reuse, the fabricated-authority attack about 45 minutes and the prefilled-approval attack about 150 minutes. The unmodified transfer of a known jailbreak across a model generation shows the underlying weakness went unpatched, and the researchers framed the gap between direct-request refusals and adversarial performance as a case study in the limits of surface-level safety testing for open-weight releases.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026