Terminal-Bench
AITerminal-Bench: A benchmark of tasks an agent must complete in a real shell, testing tool use and recovery rather than code generation alone.
Used in these stories
- Meta ships Muse Code, and its own charts put Claude Code first
- Moonshot opens Kimi K3's weights: what the licence allows, and what hosting 2.8 trillion parameters really costs
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch
- GPT-5.6 Sol, Terra and Luna reset OpenAI pricing
- Grok 4.5 buys its way into the coding race
- Claude Sonnet 5 puts Opus class agents on a mid tier budget