YFarmX logoYFarmX

Tools Model Security Capability gpt-daybreak-blue

Model record

GPT Daybreak Blue: how it does on security work, and what it will answer

Lab
OpenAI
Released
2026-08-10
Access
gated
Weights
closed
Safeguards
Trusted Access for Cyber, defensive tier: identity verification, legal attestation, account monitoring
Score rows
1
Confidence
CONFIRMED

What happened

The highest-scoring model row on RealVuln: F3 79.5 across all 140 repositories at 63.98% precision and 81.75% recall, $3.91 a repository, in 347.2 seconds a run. It is GPT-5.6 Sol's weights served to verified defenders under the Daybreak programme, and the 4.8-point gap to Sol on the same board is consistent with the permitted tier spending fewer turns declining.

RealVuln 3.1.0, F3 (micro): 79.5 (independent; Codex CLI (gpt-daybreak-blue-codex-cli), prompt sha256:45a1200d61e6 (tsjs-v1); 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: Daybreak Blue tier: defensive tasks on mainline weights; run 2026-09-11; CONFIRMED). strict F3 79.5; precision 63.98%, recall 81.75%; $547.56 for the run; $0.0739 per 100 lines; 347.2s wall clock a run; reasoning effort high. Refusal gate (SINGLE, observed 2026-09-11): Defensive tasks on mainline weights for verified organisations; Blue approval does not carry into Red.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026