SWE-bench
AISWE-bench: A benchmark that gives a model real GitHub issues from open-source projects and checks whether its code patch makes the repository's tests pass.
Used in these stories
- Meta opens Muse Glimmer under Apache 2.0, dropping the Llama licence rules
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch
- GPT-5.6 Sol, Terra and Luna reset OpenAI pricing
- Grok 4.5 buys its way into the coding race
- LongCat-2.0 trained without a single Nvidia chip
- Claude Sonnet 5 puts Opus class agents on a mid tier budget