What GLM-5.3-Flash Actually Is
Strip away the stealth-test intrigue and GLM-5.3-Flash is, on paper, a fairly conventional frontier-tier mixture-of-experts design taken to an efficient extreme: 320 billion total parameters, but only 18 billion active per token [1]. That sparsity is the entire economic story of the model - a fraction of the compute of a dense 320B system gets routed per request, which is what makes the aggressive pricing possible in the first place. It's also the first natively multimodal entry in the GLM-5 line, built to handle text, images, video, and visual documents inside a single 1,048,576-token context window [2].
The pricing Zhipu attached to that architecture is deliberately undercutting: $0.15 per million input tokens, $0.03 for cached input, and $0.50 per million output tokens, with a 50% launch discount running through September 9, 2026 [2]. Technical walkthroughs circulating after launch pointed to the same efficiency logic - 18 billion active parameters out of 320 billion total - as the direct explanation for why Zhipu can serve a frontier-scale model this cheaply without the compute footprint a dense model of similar quality would require.


