Z.ai confirmed on 27 August that Ox Alpha, an anonymous model that had led OpenRouter’s charts for a week, is its own GLM-5.3-Flash. Released on 20 August under the codename, it processed more than 11 trillion tokens in its first three days, OpenRouter’s largest launch on record, and by Thursday it made up nearly 31% of coding traffic there. Its Hong Kong listed shares rose more than 8%.
The claim drawing attention is about hardware, not benchmarks. Z.ai says every online request during the stealth run was served from a cluster of 100,000 domestically produced accelerators, with no Nvidia parts in the inference path. GLM-5.3-Flash is a cheaper version of the flagship, on the same roughly 700 billion parameter base as the June model, and it sits tenth on Artificial Analysis’s intelligence index, ahead of DeepSeek’s V4 Pro Max. The weights went out on Wednesday.
For a Chinese lab working under US export controls, showing that frontier adjacent inference runs at volume on local silicon is the bigger demonstration. It tells domestic customers they can commit to Chinese cloud capacity without a stock of restricted chips, and Washington that the controls are shaping architecture, not stopping it. Zhipu, its parent, is preparing a dual listing and has grown as Anthropic pulled back from China.
What Z.ai will not supply is the detail that would make the claim checkable. It has declined to name the chips, and CNBC said it could not verify that the cluster ran without Nvidia parts. A cheap, openly weighted model is simple to confirm. A serving stack that owes nothing to American hardware is a far larger claim, and for now rests on the word of a company whose shares have already moved on it.



