<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>YFarmX Model Security Capability</title><description>Newly scored models on the YFarmX Model Security Capability tracker: how each frontier model does on security work, under which harness and safeguard state, and what it will answer.</description><link>https://yfarmx.com/</link><language>en-GB</language><item><title>Muse Spark: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/muse-spark/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/muse-spark/</guid><description>Muse Spark scores 65.4% end-to-end on Cybench on all 40 tasks, the highest full-task-set card figure on the leaderboard read 11 September 2026.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Meta</category><category>api</category><category>CONFIRMED</category></item><item><title>Claude Opus 4.5: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/claude-opus-4-5/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/claude-opus-4-5/</guid><description>Claude Opus 4.5 scores 82% end-to-end on Cybench on the 39 of 40 tasks Anthropic ran, per the benchmark&apos;s own leaderboard read on 11 September 2026.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>api</category><category>CONFIRMED</category></item><item><title>Claude Opus 4.6: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/claude-opus-4-6/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/claude-opus-4-6/</guid><description>Claude Opus 4.6 scores 93% end-to-end on Cybench on the 37 of 40 tasks Anthropic ran, per the benchmark&apos;s own leaderboard read on 11 September 2026.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>api</category><category>CONFIRMED</category></item><item><title>Claude Opus 4.7: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/claude-opus-4-7/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/claude-opus-4-7/</guid><description>Claude Opus 4.7 scores 96% end-to-end on Cybench on the 35 of 40 tasks Anthropic ran, per the benchmark&apos;s own leaderboard read on 11 September 2026.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>api</category><category>CONFIRMED</category></item><item><title>Claude Mythos Preview: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/claude-mythos-preview/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/claude-mythos-preview/</guid><description>Claude Mythos Preview scores 100% end-to-end on Cybench on the 35 of 40 tasks Anthropic ran, per the benchmark&apos;s own leaderboard read on 11 September 2026. The task count is the honest way to state it: saturated on the subset the vendor ran.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>api</category><category>CONFIRMED</category></item><item><title>Gemma 4 31B: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gemma-4-31b/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gemma-4-31b/</guid><description>The highest precision on the RealVuln board, 90.12%, at 23.50% recall and F3 25.4 on the Python subset, run on local hardware and billed at zero. A first-pass screen that reports few findings and is usually right about them.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>Qwen3.8-Max: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/qwen3-8-max/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/qwen3-8-max/</guid><description>A search-derived CVE campaign puts it at 26 of 32 recent CVEs, 81.25% pass@3, for $821.35 over three runs. Its smaller sibling Qwen3.8-27B ships under Apache-2.0 and had 150 refusal-removed rebuilds within a week of release.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Alibaba</category><category>open-weights</category><category>SINGLE</category></item><item><title>CyberKimi: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/cyberkimi/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/cyberkimi/</guid><description>Kimi K3 with the refusal layer removed and cyber post-training added. On the same V8 bug, harness and prompt as stock Kimi K3 it scored 8 of 16 capabilities unassisted and 10 of 16 with a methodology pack against the stock model&apos;s 4, the cleanest published control on how much willingness moves a security score.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Adverserial AI</category><category>waitlist</category><category>CONFIRMED</category></item><item><title>Kimi K3: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/kimi-k3/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/kimi-k3/</guid><description>The value pick on RealVuln: F3 59.4 on the Python subset at 73.02% precision for $21.33, $0.32 a repository. DeepSeek&apos;s table gives it 80.0 on CyberGym, and on ExploitBench against CVE-2024-6100 the stock model scored 4 of 16 capabilities. It is the model credited with a working Redis exploit, &quot;the first llm that is capable and willing to write an exploit&quot;.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Moonshot AI</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>DeepSeek V4 Pro: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/deepseek-v4-pro/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/deepseek-v4-pro/</guid><description>RealVuln scores it F3 38.0 across all 140 repositories at 60.91% precision over three runs; a search-derived CVE campaign puts it at 87.5% pass@3 at 65.6% precision on 32 recent CVEs.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>DeepSeek</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>DeepSeek V4 Flash: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/deepseek-v4-flash/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/deepseek-v4-flash/</guid><description>The cheapest measured pass on RealVuln: all 140 repositories at $0.0005 per 100 lines, F3 41.8 at 56.31% precision, $3.83 in total over three trials. A 120,000-line project costs about $0.62.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>DeepSeek</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>DeepSeek V4.1 Flash: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/deepseek-v4-1-flash/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/deepseek-v4-1-flash/</guid><description>Released 10 September 2026 under MIT. DeepSeek&apos;s own table gives 88.1 on CyberGym; RealVuln has it at F3 50.9 across all 140 repositories at 97.8 seconds a run, the fastest full-corpus wall clock on the board, billed at zero because it ran on local hardware.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>DeepSeek</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>GLM-5.3: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/glm-5-3/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/glm-5-3/</guid><description>Z.ai&apos;s own table gives 84.5 on CyberGym and 54.4 on ExploitBench; RealVuln has it at F3 56.9 across 136 of 140 repositories at $0.0148 per 100 lines, with four validation failures. It is the default model in the Strix pentesting agent&apos;s quickstart.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Z.ai</category><category>open-weights</category><category>CONFIRMED</category></item><item><title>Grok 4.6: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/grok-4-6/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/grok-4-6/</guid><description>On 28 August 2026 the Shannon agent ran Grok 4.6 through an authorised pentest of Photoview 2.4.0: 10 reported findings, 10 true positives, no false positives, for $35.07 in 5 hours 26 minutes, with the SARIF output opened by this desk. xAI is the one lab that publishes the text it injects in front of the model.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>xAI</category><category>api</category><category>CONFIRMED</category></item><item><title>Gemini 3.5 Flash: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gemini-3-5-flash/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gemini-3-5-flash/</guid><description>The precision-shaped row on RealVuln: 89.77% precision at 33.26% recall, F3 35.5, on 64 of 140 repositories over three runs, with one timeout and one validation failure. It reports few findings and is usually right about them.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>api</category><category>CONFIRMED</category></item><item><title>Gemini 3.5 Flash Cyber: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gemini-3-5-flash-cyber/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gemini-3-5-flash-cyber/</guid><description>The first Flash Cyber model, credited by Google with 55 unique confirmed V8 issues on Big Sleep; its CyberGym figure is contested between 83.2% and 77.5% across Google&apos;s own materials.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>gated</category><category>CONFIRMED</category></item><item><title>Gemini 3.8 Flash Cyber: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gemini-3-8-flash-cyber/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gemini-3-8-flash-cyber/</guid><description>The model inside CodeMender since 2 September 2026, reachable only through Fairwind. Google&apos;s own table gives 86.2% pass@1 on CyberGym, 47.2% on CWE-Bench at roughly $3.60 a rollout, and a 6.0% attack success rate on Gray Swan injection.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>gated</category><category>CONFIRMED</category></item><item><title>GPT-5.5: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gpt-5-5/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gpt-5-5/</guid><description>RealVuln scores GPT-5.5 at F3 56.7 on the Python subset at 72.62% precision over three runs, and Google&apos;s comparison table quotes GPT-5.5-Cyber at 85.6% pass@1 on CyberGym.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>api</category><category>CONFIRMED</category></item><item><title>GPT Daybreak Red: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gpt-daybreak-red/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gpt-daybreak-red/</guid><description>The only model in the field sold for &quot;proof-of-concept exploit development, exploit-chain validation, penetration testing, and red teaming&quot;, at $12.50 and $75 per million tokens with a 400,000-token window. No public benchmark row exists for it.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>gated</category><category>SINGLE</category></item><item><title>GPT Daybreak Blue: how it does on security work, and what it will answer</title><link>https://yfarmx.com/tools/model-security-capability/gpt-daybreak-blue/</link><guid isPermaLink="true">https://yfarmx.com/tools/model-security-capability/gpt-daybreak-blue/</guid><description>The highest-scoring model row on RealVuln: F3 79.5 across all 140 repositories at 63.98% precision and 81.75% recall, $3.91 a repository, in 347.2 seconds a run. It is GPT-5.6 Sol&apos;s weights served to verified defenders under the Daybreak programme, and the 4.8-point gap to Sol on the same board is consistent with the permitted tier spending fewer turns declining.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>gated</category><category>CONFIRMED</category></item></channel></rss>