Z.ai / Zhipu
GLM-5.3
a post-training update with an emergent cyber capability

Key facts
- 14 Aug 2026post-training only
- Released
- 743Bper Z.ai, unchanged from GLM-5.2
- Base model
- 1M / 128ktokens in / out
- Context
- 28 Aug 2026custom glm-5.3 licence
- Weights
- coming soonCoding Plan only for now
- API
- 84.5%best on Z.ai's own table
- CyberGym
Z.ai took the GLM-5.2 base and post-trained it harder. Coding improved. What jumped was the ability to find and chain software exploits, which is why the weights were held back for a fortnight of safety hardening. They arrived on 28 August 2026, under a bespoke glm-5.3 licence rather than GLM-5.2's MIT.
What it is
GLM-5.3 is the flagship model from the Chinese lab Z.ai, formerly Zhipu, announced on 14 August 2026. It runs the GLM-5.2 base with every gain bought by scaled-up post-training. Z.ai states it directly, in the blog post and again in its own documentation: “it uses the same base model as GLM-5.2, and every gain comes from post-training.”
Post-training is everything done to a model after the expensive pre-training run that builds its raw knowledge: reinforcement learning, tuning against tasks, teaching it to use tools. Z.ai’s summary of the release fits in one line of its own: “Scaling post-training is all we did for GLM-5.3.”
The base carries over unchanged: a mixture-of-experts model that Z.ai’s own launch post describes as a 743-billion-parameter base, with a one-million-token context window and output up to 128,000 tokens. One breaking change arrives with it. Where GLM-5.2 let a developer switch thinking off, GLM-5.3 does not: the documentation says “disabling thinking is no longer supported”, leaving three levels of reasoning effort, low, high and max.
What the extra post-training bought
On Z.ai’s own table, the coding gains over GLM-5.2 are large. Terminal-Bench 3.0 goes from 4.6 to 28.3. DeepSWE v1.1 goes from 46.2 to 66.9. Z.ai also claims better token economy: at maximum effort the new model reaches a higher score on its in-house benchmark while spending roughly 75,000 output tokens per task against GLM-5.2’s 96,000.
The cyber numbers are the ones the company leads with, and they moved further. ExploitBench, which measures writing working exploits, went from 24.4 to 54.4 per cent. ExploitGym went from 29 solved tasks to 105 at a two-hour budget. CyberGym, on vulnerability discovery, reached 84.5 per cent, the best score in Z.ai’s table.
Z.ai describes this as something it did not fully plan for:
“What surprised us was how quickly the capability continued to develop as training scaled. GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.”
Z.ai, 14 August 2026
Where it loses, by Z.ai’s own count
The launch table compares GLM-5.3 against seven rivals across sixteen rows, and Z.ai’s model takes the top score on three: CyberGym, AutomationBench and GDPVal-AA v2. The rest go elsewhere, and several are not close. Nine of the sixteen rows are below, with the leader named in the last column; the full table is on Z.ai’s own page.
| Benchmark | GLM-5.3 | GLM-5.2 | Leader |
|---|---|---|---|
| Terminal Bench 3.0 | 28.3 | 4.6 | 34.6 GPT-5.6 Sol |
| DeepSWE v1.1 | 66.9 | 46.2 | 72.7 GPT-5.6 Sol |
| FrontierSWE | 78.1 | 67.5 | 88.2 Fable 5 |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.8 Opus 4.8 |
| CyberGym | 84.5 | 77.2 | GLM-5.3 |
| ExploitBench | 54.4 | 24.4 | 78.0 Fable 5 |
| Toolathlon Verified | 73.0 | 59.9 | 76.5 Kimi K3 |
| AutomationBench | 48.2 | 26.2 | GLM-5.3 |
| GDPVal-AA v2 | 1769 | 1508 | GLM-5.3 |
On the cyber benchmarks the company leads with, the closed models are still well clear: Claude Fable 5 scores 78.0 on ExploitBench against GLM-5.3’s 54.4. Z.ai says so itself, and the sentence is worth reading twice:
“The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2, and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.”
Z.ai, 14 August 2026
Two things about how these figures were produced deserve recording. Most of the agentic and cyber rows were run inside Claude Code 2.1.207, so Z.ai tested its own model and its rivals using Anthropic’s coding agent as the rig. And the one row not run on Z.ai’s own harness, GDPVal-AA v2, was measured by Artificial Analysis, which is also the row with the widest margin in Z.ai’s favour.
One oddity in the source: the table column reads “Fable 5 (w/ fallback)” while the body text calls the same figures “Mythos 5”. Those are two real Anthropic models, released the same day, and our own reference pages describe Mythos 5 as the restricted twin of Fable 5, so the confusion is understandable. It is still Z.ai’s own document disagreeing with itself, and the numbers quoted in the prose match the Fable 5 column.
Access: the Coding Plan first, the weights later
This was the detail most coverage got wrong at launch. GLM-5.3 was not generally available, and Z.ai’s documentation said so in a banner at the top of the model page: “The GLM-5.3 API is coming soon.”
What existed at launch was access through the GLM Coding Plan and Z.ai’s ZCode agent, with plans listed at $12.60, $56 and $117.60 a month, and a points system that charges half rate outside 14:00 to 18:00 China time on weekdays. At that point there was no GLM-5.3 row on Z.ai’s pricing table, no repository on Hugging Face, no listing on OpenRouter, and no Artificial Analysis score, because independent scoring generally needs an API to test against.
The weights promise, “we will release the weights in two weeks after launch, once safety evaluation and hardening are complete”, was kept roughly on schedule. On 28 August 2026 the full model landed in Z.ai’s Hugging Face organisation as zai-org/GLM-5.3: 141 safetensors shards totalling 756 GB, with a BF16 variant published alongside. The licence is the change to note. GLM-5.2 shipped under MIT within three days of its API; GLM-5.3 carries a bespoke glm-5.3 licence instead, so for this release open weight is the accurate description, not open source. The model card lists 753 billion parameters, ten billion more than the launch post’s 743-billion figure for the shared base, and repeats the CyberGym score of 84.5 against GLM-5.2’s 77.2.
The fortnight also produced a sibling. On 26 August 2026 Z.ai introduced GLM-5.3-Flash, a smaller open release at 320 billion total parameters with 18 billion active, natively multimodal, with a one-million-token context and an MIT licence, and said it was the model it had been testing anonymously on OpenRouter as Ox Alpha since 20 August.
For the model underneath this one, see GLM-5.2. For the open-weight rivals in the table, see Kimi K3 and DeepSeek V4 Pro.
More in Large Language Models
All LLMs →- Z.ai / ZhipuGLM-5.2strongest open-weight all-rounder before K3 landed
- Moonshot AIKimi K3the largest open-weight model yet released
- DeepSeekDeepSeek V4 Prothe price disruptor, raising its prices
- OpenAIGPT-5.6the flagship from 9 July to 3 September 2026
- OpenAIGPT-6 Astrathe new generation, rolling out from 3 September 2026
- AnthropicClaude Fable 5the flagship holding first place on the Arena text board
