AI News

Ox Alpha was GLM-5.3-Flash, a model that did not exist publicly when we fingerprinted it

Z.ai confirmed on 26 August that the anonymous OpenRouter model Ox Alpha was GLM-5.3-Flash. Our 21 August fingerprint put it on the GLM 5 line and flagged one thing that did not fit. That detail is what the answer turned on.

Editorial hero: a halftone ox photographed in profile at large scale on clear newsprint, a paper luggage tag at its neck reading GLM-5.3-FLASH, headed OX ALPHA with the subtitle CONFIRMED AS GLM-5.3-FLASH

On 21 August we published a fingerprint of Ox Alpha, the anonymous model that had appeared on OpenRouter the day before with no owner attached. Two tests, one on its published API contract and one on how it counted tokens, both pointed at Z.ai’s GLM 5 line. We put it at roughly ninety-five per cent and named the thing that stopped us going higher:

One detail still does not fit. Ox Alpha takes video, and Z.ai’s documentation says GLM-5.3 “currently supports text-only inputs”.

Z.ai answered on 26 August, in the documentation for a model that had not existed publicly when we ran the tests.

The confirmation

The GLM-5.3-Flash launch page says it directly:

Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback.

Z.ai’s founder Jie Tang put it more bluntly on X the next morning: “Ox Alpha = GLM-5.3 Flash”.

So the tokenizer evidence was right and the API-contract evidence was right, and the thing that did not fit was not an error in the method. It was a model we could not have named, because Z.ai had not published it. GLM-5.3-Flash is described by Z.ai as “the first natively multimodal model in the GLM-5 series”. A sibling that takes video, sharing the GLM 5 tokenizer and inheriting the GLM-5.3 API record, is exactly what our two tests would have produced.

The five per cent we held back was the right five per cent, and it was pointing at the answer rather than away from it.

What the model is

GLM-5.3-Flash is 320 billion parameters with 18 billion active, released under the MIT licence with the weights on Hugging Face the same day. Z.ai lists it at $0.15 per million input tokens and $0.50 output, currently halved to $0.075 and $0.25 under a discount its own pricing page says “ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)”.

The weights arrived on the day of the announcement. The Hugging Face repository was created as an empty placeholder at 06:43 UTC on 25 August, and the model files were uploaded between 11:58 and 13:07 the following day.

Infographic: the stealth model named, and what Z.ai said about its own anonymity. The model is GLM-5.3-Flash, 320B total and 18B active. The licence is MIT, with weights on day one. The context is 1M in and 128K out. The stated reason for anonymity was to gather user feedback. The qualifier, added later, is that Ox Alpha was an early version. Source: the Z.ai launch post titled GLM-5.3-Flash: frontier intelligence, flash cost, 26 August 2026.

The version question nobody has answered

Three minutes after the company post, Z.ai’s lead Zixuan Li added a qualifier:

Ox Alpha was an early version of GLM-5.3-Flash. The official release delivers stronger performance and significantly better stability.

That is a material change to what the confirmation means. Anyone who benchmarked Ox Alpha during its anonymous run was benchmarking something that is not the released model, and Z.ai has published no changelog, no evaluation delta and no checkpoint dates to say how far apart they are. Our own measurements were taken on the anonymous version.

Whether the two share weights is not established, and we are not going to assert it either way.

Three documents, three positions on your prompts

The part of this worth keeping is not the identity. It is what the arrangement said about the data.

OpenRouter’s Stealth Program EULA, last updated 6 July 2026, says stealth models are offered “for purposes of collecting User Content for use in Stealth Model training and improvement”. That is the payment: the prompts are the price of free access.

OpenRouter’s own listing for Ox Alpha said the opposite, that prompts and completions are retained by the provider and are not used for training.

OpenCode, the other launch partner, advertised the same model on 20 August with “Zero Data Retention”.

Three documents disagreeing about prompts sent to Ox Alpha: the Stealth Program EULA says they were collected for stealth model training and improvement, OpenRouter's own listing said they were retained but not used for training, and OpenCode advertised zero data retention.
All three were live during the same six days.

Three statements about one model, live at the same time, and they cannot all be true. Anyone who put something sensitive through Ox Alpha during those six days should assume the EULA is the operative document, because it is the one with a signature block.

It is gone, and it is still first

Ox Alpha has been withdrawn. OpenRouter’s Stealth provider page now reads:

Stealth has no models available on OpenRouter right now. Check back later.

The model’s endpoints array is empty and the Stealth provider has dropped out of OpenRouter’s provider list. Nobody announced the withdrawal, and OpenRouter published no post about Ox Alpha at any stage, before or after the reveal, which departs from what it did for earlier stealth models.

The rankings have not caught up with any of this. Ox Alpha still sits at number one on OpenRouter’s public model rankings with 20.7 trillion tokens for the week to 29 August, up 216 per cent. The named GLM-5.3-Flash listing enters separately at number eight with 4.62 trillion. The same model holds two positions because OpenRouter kept the listings separate rather than merging them.

The claim we cannot check

Z.ai says the anonymous run was served entirely on Chinese silicon. Its launch post lists it as a bullet: “Previously previewed as Ox Alpha, running entirely on Chinese AI chips”.

No vendor is named. Neither the blog nor the documentation mentions Huawei, Ascend, Cambricon, Biren, MetaX or Moore Threads, and the strongest phrase either document offers is “domestically developed accelerators”. That is a significant claim about capability if it is true, and there is no way to verify it from outside, so we record it as Z.ai’s statement rather than as a fact.

Twenty separate providers now serve GLM-5.3-Flash on OpenRouter, most at the list rate, with one undercutting Z.ai’s own endpoint at $0.05 and $0.1667 on a reduced-precision build. The stealth run is over; the price competition it was presumably meant to seed has started.

Sources

  1. GLM-5.3-Flash model documentation (Z.ai)docs.z.ai
  2. GLM-5.3-Flash: Frontier Intelligence, Flash Cost (Z.ai blog)z.ai
  3. Z.ai's launch postx.com
  4. Zixuan Li's qualifierx.com
  5. Jie Tang confirming the identityx.com
  6. Ox Alpha on OpenRouter, with the reveal banneropenrouter.ai
  7. OpenRouter's Stealth provider pageopenrouter.ai
  8. OpenRouter Stealth Program EULAopenrouter.ai
  9. GLM-5.3-Flash on Hugging Facehuggingface.co
  10. Z.ai pricingdocs.z.ai