Tools Model Security Capability gpt-6-astra
Model record
GPT-6 Astra: how it does on security work, and what it will answer
- Lab
- OpenAI
- Released
- 2026-09-03
- Access
- api
- Weights
- closed
- Safeguards
- Classifier refusals; enterprise cyber access off by default at launch; Daybreak unavailable for Astra
- Score rows
- 3
- Confidence
- CONFIRMED
What happened
OpenAI's newest and most expensive model scores 100% on ExploitBench in its own system card, with a contamination caveat reported on the same set, and refuses 91.5% of disallowed cyber requests. On RealVuln it reads all 140 repositories at F3 52.1 for $1,040.62, $7.43 a repository, behind models costing a twentieth as much.
RealVuln 3.1.0, F3 (micro): 52.1 (independent; Codex CLI (gpt-6-astra-codex-cli), prompt sha256:45a1200d61e6; 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 52.1; precision 44.73%, recall 53.09%; $1040.62 for the run; $0.1404 per 100 lines; 574.7s wall clock a run; reasoning effort high. ExploitBench, score: 100% (vendor-card; OpenAI system card, 3 September 2026; ExploitBench as run by OpenAI, described as 41 vulnerabilities; safeguards: production safeguards state unstated in the summaries read; run 2026-09-03; SINGLE). the same card reportedly flags historical-vulnerability contamination on the set; two V8 zero-days found mid-evaluation are described as disclosed to maintainers. ExploitGym, honeypot overreach: 0% (vendor-card; OpenAI system card, 3 September 2026; ExploitGym honeypot scenario; safeguards: safeguards on; GPT-5.6 Sol without production safeguards is quoted at 48.2%; run 2026-09-03; SINGLE). Refusal gate (SINGLE, observed 2026-09-11): Refused 91.5% of disallowed cyber requests against 59% for GPT-5.6 Sol (card summary); the Codex client says Daybreak is unavailable for Astra.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- OpenAI Model Specraw.githubusercontent.com/openai/model_spec/main/model_spec.…
- ExploitBench sourceyfarmx.com/ai/llms/gpt-6-astra/
- ExploitGym sourceyfarmx.com/ai/llms/gpt-6-astra/
On YFarmX
- Reference pageyfarmx.com/ai/llms/gpt-6-astra/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026