Tools Model Security Capability grok-4-6
Model record
Grok 4.6: how it does on security work, and what it will answer
- Lab
- xAI
- Released
- 2026-08-12
- Access
- api
- Weights
- closed
- Safeguards
- Published injected safety prompt: intent test, disallowed list, written-consent carve-out for official red teams
- Score rows
- 2
- Confidence
- CONFIRMED
What happened
On 28 August 2026 the Shannon agent ran Grok 4.6 through an authorised pentest of Photoview 2.4.0: 10 reported findings, 10 true positives, no false positives, for $35.07 in 5 hours 26 minutes, with the SARIF output opened by this desk. xAI is the one lab that publishes the text it injects in front of the model.
Photoview 2.4.0 authorised pentest, true positives of findings reported: 10 of 10 (independent; Shannon agent (Keygraph), 28 August 2026, SARIF opened; Photoview 2.4.0, one application; safeguards: xAI injected safety prompt as deployed; run 2026-08-28; CONFIRMED). a single-target run, carried for its cost and its zero false positives rather than as a benchmark. Cybench, unguided, solved: 43% (vendor-card; vendor system card run; 40 of 40 tasks; safeguards: as run by the vendor; safeguard state per the card; run 2026-09-11; CONFIRMED). task count is what the vendor ran, so rows are comparable only with it beside them. Refusal gate (CONFIRMED, observed 2026-09-11): Do not answer queries that show clear intent to engage in disallowed activities; high-level answers without actionable detail to general questions such as "how to hack a website?"; "unlawfully" is the operative word. The Cybench row is the Grok 4 card figure on all 40 tasks; Grok 4.1 Thinking sits at 39% and Grok 4 Fast at 30% on the same leaderboard.
Sources
- xAI, Grok 4 safety promptraw.githubusercontent.com/xai-org/grok-prompts/main/grok_4_s…
- Photoview 2.4.0 authorised pentest sourceraw.githubusercontent.com/KeygraphHQ/shannon/main/benchmark/…
- Cybench sourceraw.githubusercontent.com/cybench/cybench.github.io/main/dat…
On YFarmX
- Reference pageyfarmx.com/ai/llms/grok-4-6/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026