The Stealth Playbook: How a Free 'Mystery Model' Became a 62-Trillion-Token Coup
On August 20, 2026, an anonymous, free-to-use model called 'Ox Alpha' quietly appeared on OpenRouter, OpenCode, Cline, and Nous Research's portal, offering a 1-million-token context window and support for text, image, and video input [1]. No lab claimed it. For about a week it ran as a stealth test, absorbing real-world traffic instead of internal benchmarking - reportedly serving more than 100 trillion tokens a day in free usage and processing 62 trillion tokens in total before any company put its name on it [2]. That is an unusual way to launch a frontier-class model: skip the keynote, let the market find you first.
The unmasking came on August 27, 2026, when Zhipu AI (Z.ai) confirmed that Ox Alpha was in fact GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model, and released its weights the same day under the MIT license [3]. The reveal validated what the stealth run had already shown: GLM-5.3-Flash became OpenRouter's biggest launch to date, processing more than 11 trillion tokens in its first three days post-reveal and capturing close to 31 percent of the platform's weekly coding-model volume - the top spot [1]. Zhipu's Hong Kong-listed shares closed more than 12 percent higher the day of the announcement [1]. The stealth period was not incidental marketing; it was the model's own adoption curve, built before anyone knew whose model it was.




