Zhipu AI officially confirmed on August 26 that the previously anonymously tested model, Ox Alpha (known as "Niulai" in the Chinese community), is indeed the newly released GLM-5.3-Flash. During its free testing period, the model accumulated over 50 trillion tokens in traffic within 5 days, setting a new platform record. Zhipu AI disclosed that all its inference services are supported by a cluster of over 100,000 domestically produced chips, and its hardware efficiency and per-token cost have reached levels comparable to mainstream NVIDIA GPUs. SemiAnalysis, a semiconductor research firm, commented on this, stating that following OpenAI's self-developed inference chip Jalapeño, NVIDIA's CUDA moat is once again being tested. GLM-5.3-Flash is priced at 0.8 yuan per million input tokens and 2.8 yuan per million output tokens, which is about 1/40 of Anthropic's Claude Opus 4.8, aiming to counter the recent wave of price increases in China's AI model market.