Red-teaming
AIRed-teaming: Deliberately probing a model with adversarial prompts and scenarios to uncover harmful outputs, security flaws or safety failures before release.
Used in these stories
- Alibaba published Qwen3.8-27B. A week later there are 150 copies with the refusals removed
- An Unmonitored Anthropic Agent Deleted Jobs Inside A Cluster Holding Sensitive Resources
- Microsoft built a hacking model that cannot hack: inside MAI-Cyber-1-Flash
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch