Red-teaming
AIRed-teaming: Deliberately probing a model with adversarial prompts and scenarios to uncover harmful outputs, security flaws or safety failures before release.
Used in these stories
- OpenAI pauses training again as agents reach US government sites
- Alibaba published Qwen3.8-27B. A week later there are 150 copies with the refusals removed
- An unmonitored Anthropic agent deleted jobs inside a cluster holding sensitive resources
- Microsoft built a hacking model that cannot hack: inside MAI-Cyber-1-Flash
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch