Tools Model Security Capability gpt-daybreak-blue
Model record
GPT Daybreak Blue: how it does on security work, and what it will answer
- Lab
- OpenAI
- Released
- 2026-08-10
- Access
- gated
- Weights
- closed
- Safeguards
- Trusted Access for Cyber, defensive tier: identity verification, legal attestation, account monitoring
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
The highest-scoring model row on RealVuln: F3 79.5 across all 140 repositories at 63.98% precision and 81.75% recall, $3.91 a repository, in 347.2 seconds a run. It is GPT-5.6 Sol's weights served to verified defenders under the Daybreak programme, and the 4.8-point gap to Sol on the same board is consistent with the permitted tier spending fewer turns declining.
RealVuln 3.1.0, F3 (micro): 79.5 (independent; Codex CLI (gpt-daybreak-blue-codex-cli), prompt sha256:45a1200d61e6 (tsjs-v1); 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: Daybreak Blue tier: defensive tasks on mainline weights; run 2026-09-11; CONFIRMED). strict F3 79.5; precision 63.98%, recall 81.75%; $547.56 for the run; $0.0739 per 100 lines; 347.2s wall clock a run; reasoning effort high. Refusal gate (SINGLE, observed 2026-09-11): Defensive tasks on mainline weights for verified organisations; Blue approval does not carry into Red.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- OpenAI python client, model union naming gpt-daybreak-blue-latestraw.githubusercontent.com/openai/openai-python/main/src/open…
- Refusal policy or licenceraw.githubusercontent.com/openai/codex/main/codex-rs/tui/src…
On YFarmX
- Reference pageyfarmx.com/ai/llms/gpt-5-6/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026