YFarmX logoYFarmX

Tools Model Security Capability glm-5-3

Model record

GLM-5.3: how it does on security work, and what it will answer

Lab
Z.ai
Released
2026-08-14
Access
open-weights
Weights
open
Safeguards
Weights licence carries no cyber restriction; hosted terms prohibit generating malicious code
Score rows
3
Confidence
CONFIRMED

What happened

Z.ai's own table gives 84.5 on CyberGym and 54.4 on ExploitBench; RealVuln has it at F3 56.9 across 136 of 140 repositories at $0.0148 per 100 lines, with four validation failures. It is the default model in the Strix pentesting agent's quickstart.

RealVuln 3.1.0, F3 (micro): 56.9 (independent; agentic harness (glm-5.3-agentic-v1), prompt sha256:45a1200d61e6; 136 of 140 repositories pinned by commit SHA (Python subset); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 54.6; precision 44.89%, recall 58.7%; $105.73 for the run; $0.0148 per 100 lines; 518.7s wall clock a run. four validation_failed exits out of 140. CyberGym, score: 84.5% (vendor-card; Z.ai launch table, 14 August 2026; CyberGym; safeguards: open weights; no classifier; run 2026-08-14; CONFIRMED). ExploitBench, score: 54.4 (vendor-card; Z.ai launch table, 14 August 2026; V8 exploitation ladder as run by Z.ai; safeguards: open weights; no classifier; run 2026-08-14; CONFIRMED). Z.ai's ExploitBench scaffold is one of three the record conflates under the name. Refusal gate (CONFIRMED, observed 2026-09-11): Licence over the weights carries no use restriction on cyber work; the hosted API's terms are separate.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026