YFarmX logoYFarmX

Tools Model Security Capability cyberkimi

Model record

CyberKimi: how it does on security work, and what it will answer

Lab
Adverserial AI
Access
waitlist
Weights
closed
Safeguards
Refusal layer ablated by design, cyber post-training added
Score rows
2
Confidence
CONFIRMED

What happened

Kimi K3 with the refusal layer removed and cyber post-training added. On the same V8 bug, harness and prompt as stock Kimi K3 it scored 8 of 16 capabilities unassisted and 10 of 16 with a methodology pack against the stock model's 4, the cleanest published control on how much willingness moves a security score.

ExploitBench bench-v8, capabilities reached on CVE-2024-6100: 8 of 16 (vendor-card; Adverserial AI campaign, 8 and 9 August 2026, unassisted; one V8 type-confusion bug, 400-turn episodes; safeguards: refusal layer ablated; run 2026-08-09; CONFIRMED). 10 of 16 with the methodology pack. CyberGym, first pass: 65.6% (vendor-card; Adverserial AI campaign; CyberGym; safeguards: refusal layer ablated; run 2026-08-09; CONFIRMED). 86.7% as a union over passes. Refusal gate (CONFIRMED, observed 2026-09-11): Ablated by design; hosted only, behind a waitlist.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026