Z.ai / Zhipu

GLM-5.3

a post-training update with an emergent cyber capability

Released 14 August 20266 min readLarge Language ModelsLast updated:

Editorial collage headed GLM-5.3 with the subtitle same base, post-training only, weights held: the Z.ai wordmark on a taped white card, a server blade labelled GLM-5.2 base unchanged, a brass dial reading post-training, a torn card showing ExploitBench 24.4 per cent rising to 54.4 per cent, a padlocked crate stencilled weights two weeks, and a torn strip reading the GLM-5.3 API is coming soon

Key facts

14 Aug 2026post-training only
Released
743Bper Z.ai, unchanged from GLM-5.2
Base model
1M / 128ktokens in / out
Context
28 Aug 2026custom glm-5.3 licence
Weights
coming soonCoding Plan only for now
API
84.5%best on Z.ai's own table
CyberGym

Z.ai took the GLM-5.2 base and post-trained it harder. Coding improved. What jumped was the ability to find and chain software exploits, which is why the weights were held back for a fortnight of safety hardening. They arrived on 28 August 2026, under a bespoke glm-5.3 licence rather than GLM-5.2's MIT.

What it is

GLM-5.3 is the flagship model from the Chinese lab Z.ai, formerly Zhipu, announced on 14 August 2026. It runs the GLM-5.2 base with every gain bought by scaled-up post-training. Z.ai states it directly, in the blog post and again in its own documentation: “it uses the same base model as GLM-5.2, and every gain comes from post-training.”

Post-training is everything done to a model after the expensive pre-training run that builds its raw knowledge: reinforcement learning, tuning against tasks, teaching it to use tools. Z.ai’s summary of the release fits in one line of its own: “Scaling post-training is all we did for GLM-5.3.”

Screenshot of Z.ai's announcement headed GLM-5.3: Frontier Coding with Emergent Cyber Capabilities, dated 2026-08-14, with buttons for the coding plan, the technical blog and a Hugging Face link marked Coming Soon
Z.ai's announcement, 14 August 2026. Note the Hugging Face button: "Coming Soon". Screenshot of z.ai/blog/glm-5.3, 14 August 2026.

The base carries over unchanged: a mixture-of-experts model that Z.ai’s own launch post describes as a 743-billion-parameter base, with a one-million-token context window and output up to 128,000 tokens. One breaking change arrives with it. Where GLM-5.2 let a developer switch thinking off, GLM-5.3 does not: the documentation says “disabling thinking is no longer supported”, leaving three levels of reasoning effort, low, high and max.

What the extra post-training bought

On Z.ai’s own table, the coding gains over GLM-5.2 are large. Terminal-Bench 3.0 goes from 4.6 to 28.3. DeepSWE v1.1 goes from 46.2 to 66.9. Z.ai also claims better token economy: at maximum effort the new model reaches a higher score on its in-house benchmark while spending roughly 75,000 output tokens per task against GLM-5.2’s 96,000.

The cyber numbers are the ones the company leads with, and they moved further. ExploitBench, which measures writing working exploits, went from 24.4 to 54.4 per cent. ExploitGym went from 29 solved tasks to 105 at a two-hour budget. CyberGym, on vulnerability discovery, reached 84.5 per cent, the best score in Z.ai’s table.

Z.ai describes this as something it did not fully plan for:

“What surprised us was how quickly the capability continued to develop as training scaled. GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.”

Z.ai, 14 August 2026

One base model, more post-training, and a capability the company says grew faster than expected.

Where it loses, by Z.ai’s own count

The launch table compares GLM-5.3 against seven rivals across sixteen rows, and Z.ai’s model takes the top score on three: CyberGym, AutomationBench and GDPVal-AA v2. The rest go elsewhere, and several are not close. Nine of the sixteen rows are below, with the leader named in the last column; the full table is on Z.ai’s own page.

Benchmark GLM-5.3 GLM-5.2 Leader
Terminal Bench 3.0 28.3 4.6 34.6 GPT-5.6 Sol
DeepSWE v1.1 66.9 46.2 72.7 GPT-5.6 Sol
FrontierSWE 78.1 67.5 88.2 Fable 5
SWE-Marathon v1.1 42.5 19.4 48.8 Opus 4.8
CyberGym 84.5 77.2 GLM-5.3
ExploitBench 54.4 24.4 78.0 Fable 5
Toolathlon Verified 73.0 59.9 76.5 Kimi K3
AutomationBench 48.2 26.2 GLM-5.3
GDPVal-AA v2 1769 1508 GLM-5.3

On the cyber benchmarks the company leads with, the closed models are still well clear: Claude Fable 5 scores 78.0 on ExploitBench against GLM-5.3’s 54.4. Z.ai says so itself, and the sentence is worth reading twice:

“The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2, and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.”

Z.ai, 14 August 2026

Two things about how these figures were produced deserve recording. Most of the agentic and cyber rows were run inside Claude Code 2.1.207, so Z.ai tested its own model and its rivals using Anthropic’s coding agent as the rig. And the one row not run on Z.ai’s own harness, GDPVal-AA v2, was measured by Artificial Analysis, which is also the row with the widest margin in Z.ai’s favour.

One oddity in the source: the table column reads “Fable 5 (w/ fallback)” while the body text calls the same figures “Mythos 5”. Those are two real Anthropic models, released the same day, and our own reference pages describe Mythos 5 as the restricted twin of Fable 5, so the confusion is understandable. It is still Z.ai’s own document disagreeing with itself, and the numbers quoted in the prose match the Fable 5 column.

Access: the Coding Plan first, the weights later

This was the detail most coverage got wrong at launch. GLM-5.3 was not generally available, and Z.ai’s documentation said so in a banner at the top of the model page: “The GLM-5.3 API is coming soon.”

Screenshot of Z.ai's GLM-5.3 documentation page showing a banner reading The GLM-5.3 API is coming soon, and beneath it GLM-5.3 is now available to all GLM Coding Plan users
Z.ai's own documentation, showing what is and is not open. Screenshot of docs.z.ai, 14 August 2026.

What existed at launch was access through the GLM Coding Plan and Z.ai’s ZCode agent, with plans listed at $12.60, $56 and $117.60 a month, and a points system that charges half rate outside 14:00 to 18:00 China time on weekdays. At that point there was no GLM-5.3 row on Z.ai’s pricing table, no repository on Hugging Face, no listing on OpenRouter, and no Artificial Analysis score, because independent scoring generally needs an API to test against.

The weights promise, “we will release the weights in two weeks after launch, once safety evaluation and hardening are complete”, was kept roughly on schedule. On 28 August 2026 the full model landed in Z.ai’s Hugging Face organisation as zai-org/GLM-5.3: 141 safetensors shards totalling 756 GB, with a BF16 variant published alongside. The licence is the change to note. GLM-5.2 shipped under MIT within three days of its API; GLM-5.3 carries a bespoke glm-5.3 licence instead, so for this release open weight is the accurate description, not open source. The model card lists 753 billion parameters, ten billion more than the launch post’s 743-billion figure for the shared base, and repeats the CyberGym score of 84.5 against GLM-5.2’s 77.2.

The fortnight also produced a sibling. On 26 August 2026 Z.ai introduced GLM-5.3-Flash, a smaller open release at 320 billion total parameters with 18 billion active, natively multimodal, with a one-million-token context and an MIT licence, and said it was the model it had been testing anonymously on OpenRouter as Ox Alpha since 20 August.

For the model underneath this one, see GLM-5.2. For the open-weight rivals in the table, see Kimi K3 and DeepSeek V4 Pro.