The Post-Training Ceiling Just Moved
Zhipu says GLM-5.3 runs on the exact same 743B-parameter base model as GLM-5.2 - every reported gain comes from scaled post-training, not a bigger or retrained base model [1]. Terminal-Bench 3.0 jumped from 4.6 to 28.3 and DeepSWE v1.1 climbed from 46.2 to 66.9 [1][2]. Zhipu says the post-training moved beyond isolated coding puzzles into full engineering workflows - identifying, analyzing, implementing, verifying, and delivering fixes - using real compute clusters, storage systems, and internal code libraries, with some tasks resembling multi-day senior-engineer workloads [2]. On an internal Z.ai benchmark, GLM-5.3 also outscored a comparison result (31.4% vs 29.5%) while using roughly 50,000 output tokens per task versus about 120,000 tokens for the reference run, suggesting the model got more efficient as well as more capable [1].


