Tools Model Security Capability gpt-5-6-sol
Model record
GPT-5.6 Sol: how it does on security work, and what it will answer
- Lab
- OpenAI
- Released
- 2026-07-09
- Access
- api
- Weights
- closed
- Safeguards
- Classifier refusals; a refused turn retries once on the configured Daybreak model
- Score rows
- 2
- Confidence
- CONFIRMED
What happened
$4 and $20 per million tokens. RealVuln scores it F3 74.7 across all 140 repositories at 59.33% precision and 76.92% recall, 4.8 points behind Daybreak Blue on identical weights, and OpenAI's own table gives 78.5% on ExploitBench. In Codex a refused turn retries once on the Daybreak model without changing the session's stored model.
RealVuln 3.1.0, F3 (micro): 74.7 (independent; Codex CLI (gpt-5.6-sol-codex-cli), prompt sha256:45a1200d61e6 (tsjs-v1); 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 74.7; precision 59.33%, recall 76.92%; $560.34 for the run; $0.0756 per 100 lines; 568.6s wall clock a run; reasoning effort high. ExploitBench, score: 78.5% (vendor-card; OpenAI comparison table; ExploitBench as run by OpenAI; safeguards: safeguard state unstated in the summaries read; run 2026-09-03; SINGLE). Refusal gate (CONFIRMED, observed 2026-09-11): Refused turns retry once on the configured Daybreak model; only a refused turn routes there.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- OpenAI Codex, Daybreak eligibility moduleraw.githubusercontent.com/openai/codex/main/codex-rs/tui/src…
- ExploitBench sourceyfarmx.com/ai/llms/gpt-6-astra/
On YFarmX
- Reference pageyfarmx.com/ai/llms/gpt-5-6/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026