Z.ai / Zhipu
GLM-5.3
a post-training update with an emergent cyber capability

Key facts
- 14 Aug 2026post-training only
- Released
- 743Bunchanged from GLM-5.2
- Base model
- 1M / 128ktokens in / out
- Context
- 2 weeksheld for safety hardening
- Weights
- coming soonCoding Plan only for now
- API
- 84.5%best on Z.ai's own table
- CyberGym
Z.ai took the GLM-5.2 base, changed nothing about it, and post-trained it harder. Coding improved. What jumped was the ability to find and chain software exploits, which is why the weights are being held back for a fortnight and why the company's own write-up is unusually candid about where it still trails.
What it is
GLM-5.3 is the flagship model from the Chinese lab Z.ai, formerly Zhipu, announced on 14 August 2026. The unusual part is what it is not: a new model. Z.ai states it directly, in the blog post and again in its own documentation, that “it uses the same base model as GLM-5.2, and every gain comes from post-training.”
Post-training is everything done to a model after the expensive pre-training run that builds its raw knowledge: reinforcement learning, tuning against tasks, teaching it to use tools. Z.ai’s summary of the release fits in one line of its own: “Scaling post-training is all we did for GLM-5.3.”
The base carries over unchanged: a mixture-of-experts model that Z.ai’s own launch post describes as a 743-billion-parameter base, with a one-million-token context window and output up to 128,000 tokens. One breaking change arrives with it. Where GLM-5.2 let a developer switch thinking off, GLM-5.3 does not: the documentation says “disabling thinking is no longer supported”, leaving three levels of reasoning effort, low, high and max.
What the extra post-training bought
On Z.ai’s own table, the coding gains over GLM-5.2 are large. Terminal-Bench 3.0 goes from 4.6 to 28.3. DeepSWE v1.1 goes from 46.2 to 66.9. Z.ai also claims better token economy: at maximum effort the new model reaches a higher score on its in-house benchmark while spending roughly 75,000 output tokens per task against GLM-5.2’s 96,000.
The cyber numbers are the ones the company leads with, and they moved further. ExploitBench, which measures writing working exploits, went from 24.4 to 54.4 per cent. ExploitGym went from 29 solved tasks to 105 at a two-hour budget. CyberGym, on vulnerability discovery, reached 84.5 per cent, the best score in Z.ai’s table.
Z.ai describes this as something it did not fully plan for:
“What surprised us was how quickly the capability continued to develop as training scaled. GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.”
Z.ai, 14 August 2026
Where it loses, by Z.ai’s own count
The launch table compares GLM-5.3 against seven rivals across sixteen rows, and Z.ai’s model takes the top score on three: CyberGym, AutomationBench and GDPVal-AA v2. The rest go elsewhere, and several are not close. Nine of the sixteen rows are below, with the leader named in the last column; the full table is on Z.ai’s own page.
| Benchmark | GLM-5.3 | GLM-5.2 | Leader |
|---|---|---|---|
| Terminal Bench 3.0 | 28.3 | 4.6 | 34.6 GPT-5.6 Sol |
| DeepSWE v1.1 | 66.9 | 46.2 | 72.7 GPT-5.6 Sol |
| FrontierSWE | 78.1 | 67.5 | 88.2 Fable 5 |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.8 Opus 4.8 |
| CyberGym | 84.5 | 77.2 | GLM-5.3 |
| ExploitBench | 54.4 | 24.4 | 78.0 Fable 5 |
| Toolathlon Verified | 73.0 | 59.9 | 76.5 Kimi K3 |
| AutomationBench | 48.2 | 26.2 | GLM-5.3 |
| GDPVal-AA v2 | 1769 | 1508 | GLM-5.3 |
On the cyber benchmarks the company leads with, the closed models are still well clear: Claude Fable 5 scores 78.0 on ExploitBench against GLM-5.3’s 54.4. Z.ai says so itself, and the sentence is worth reading twice:
“The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2, and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.”
Z.ai, 14 August 2026
Two things about how these figures were produced deserve recording. Most of the agentic and cyber rows were run inside Claude Code 2.1.207, so Z.ai tested its own model and its rivals using Anthropic’s coding agent as the rig. And the one row not run on Z.ai’s own harness, GDPVal-AA v2, was measured by Artificial Analysis, which is also the row with the widest margin in Z.ai’s favour.
One oddity in the source: the table column reads “Fable 5 (w/ fallback)” while the body text calls the same figures “Mythos 5”. Those are two real Anthropic models, released the same day, and our own reference pages describe Mythos 5 as the restricted twin of Fable 5, so the confusion is understandable. It is still Z.ai’s own document disagreeing with itself, and the numbers quoted in the prose match the Fable 5 column.
You cannot use it yet
This is the detail most coverage gets wrong. GLM-5.3 is not generally available, and Z.ai’s documentation says so in a banner at the top of the model page: “The GLM-5.3 API is coming soon.”
What exists today is access through the GLM Coding Plan and Z.ai’s ZCode agent, with plans listed at $12.60, $56 and $117.60 a month, and a points system that charges half rate outside 14:00 to 18:00 China time on weekdays. There is no GLM-5.3 row on Z.ai’s pricing table, no repository on Hugging Face, no listing on OpenRouter, and no Artificial Analysis score, because independent scoring generally needs an API to test against.
The weights are promised, with a condition attached: “we will release the weights in two weeks after launch, once safety evaluation and hardening are complete.” Z.ai has not said which licence they will carry. GLM-5.2 shipped under MIT within three days of its API, so the precedent is permissive and quick; this release is neither, and the company ties the delay to the same cyber capability it is advertising.
For the model underneath this one, see GLM-5.2. For the open-weight rivals in the table, see Kimi K3 and DeepSeek V4 Pro.
More in Large Language Models
All LLMs →- OpenAIGPT-5.6the flagship since 9 July 2026
- AnthropicClaude Fable 5the flagship topping most July 2026 rankings
- AnthropicClaude Opus 5frontier work with a dial on the bill
- AnthropicClaude Mythos 5restricted twin of Fable 5
- AnthropicClaude Sonnet 5the speed and intelligence balance
- Google DeepMindGemini 3.5 familythe generation behind Gemini today