Terminal-Bench
AITerminal-Bench: A benchmark of tasks an agent must complete in a real shell, testing tool use and recovery rather than code generation alone.
Used in these stories
- Meta Ships Muse Code, and Its Own Charts Put Claude Code First
- Moonshot opens Kimi K3's weights: what the licence allows, and what hosting 2.8 trillion parameters really costs
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch
- GPT-5.6 Sol Terra And Luna Reset OpenAI Pricing
- Grok 4.5 Buys Its Way Into The Coding Race
- Claude Sonnet 5 Puts Opus Class Agents On A Mid Tier Budget