Zhipu AI releases GLM-5.3 coding and cybersecurity model
TECH

Zhipu AI releases GLM-5.3 coding and cybersecurity model

29+
Signals

Strategic Overview

  • 01.
    Zhipu AI released GLM-5.3 on August 14, 2026, a post-trained update to the same 743B-parameter base model used in GLM-5.2, with all reported capability gains coming from extended post-training rather than a new or retrained base model.
  • 02.
    GLM-5.3 is available now through the GLM Coding Plan and integrates with coding agents including ZCode, Claude Code, and OpenCode, but open model weights won't publish until roughly two weeks after launch, pending security evaluation and hardening.
  • 03.
    GLM-5.3 posted large benchmark jumps over GLM-5.2, including Terminal-Bench 3.0 (4.6 to 28.3) and DeepSWE v1.1 (46.2 to 66.9), alongside gains on cybersecurity benchmarks CyberGym and ExploitBench.
  • 04.
    Working with Chinese security teams, GLM-5.3 helped surface 2,436 vulnerabilities across 269 open-source projects, with about half rated medium severity or higher, including one flaw in code written roughly 40 years ago.

The Post-Training Ceiling Just Moved

Zhipu says GLM-5.3 runs on the exact same 743B-parameter base model as GLM-5.2 - every reported gain comes from scaled post-training, not a bigger or retrained base model [1]. Terminal-Bench 3.0 jumped from 4.6 to 28.3 and DeepSWE v1.1 climbed from 46.2 to 66.9 [1][2]. Zhipu says the post-training moved beyond isolated coding puzzles into full engineering workflows - identifying, analyzing, implementing, verifying, and delivering fixes - using real compute clusters, storage systems, and internal code libraries, with some tasks resembling multi-day senior-engineer workloads [2]. On an internal Z.ai benchmark, GLM-5.3 also outscored a comparison result (31.4% vs 29.5%) while using roughly 50,000 output tokens per task versus about 120,000 tokens for the reference run, suggesting the model got more efficient as well as more capable [1].

An Unplanned Capability: Exploit Chains Zhipu Didn't Plan For

The more consequential detail isn't a benchmark line - it's that GLM-5.3 wasn't trained to do this. During cybersecurity training, the model began reasoning across multiple stages of exploitation, forming coherent plans for complete exploit chains, something Zhipu itself describes as an emergent capability [4]. That behavior is the stated reason open weights are being held back roughly two weeks past the hosted launch for safety evaluation and hardening, rather than shipping immediately [1][3]. Paired with Chinese security teams, the model was turned loose on real open-source codebases and surfaced 2,436 vulnerabilities across 269 projects, with roughly half rated medium severity or higher and logged to a public registry [3][5]. One of the flaws it caught sat in code written roughly 40 years ago [3]. Zhipu frames the dual-use tension as a feature rather than a bug, casting the release in explicitly cooperative terms: "AI development should not be a solo performance by one nation, but a symphony of global collaboration." [5]

"Strongest Open-Weights Coding Model" - Says Who?

Zhipu calls GLM-5.3 the strongest open-weights coding model available, and the cybersecurity numbers back that up in places: on CyberGym it edges out the Anthropic model referenced in coverage as Mythos 5 (84.5% vs 83.8%) and GPT-5.6 Sol (83.6%) [2][3]. But the gap reopens on ExploitBench, where GLM-5.3 climbed from 24.4% to 54.4% while Mythos 5 and GPT-5.6 Sol both scored in the mid-70s to high-70s [1][2]- a wide enough margin to complicate the 'strongest' framing for offensive security specifically. Community reviewers pushed on the coding claim too, pointing out that in Zhipu's own published benchmark table, another already-downloadable open-weights model scores higher than GLM-5.3 on four separate rows, a nuance the official announcement didn't mention. Pricing drew similar skepticism: threads on the GLM Coding Plan called it overpriced next to other open models, with recurring worry that each GLM version bump quietly raises thinking-token usage - and therefore real-world cost - even as headline benchmarks improve. Roleplay-focused users were unmoved altogether, noting the gains are coding-specific and don't carry into general chat quality.

Delayed Weights, Nervous Market: Security Review or PR Cover?

The market wasn't fully convinced by the security-first pitch either: Z.ai's Hong Kong-listed shares slid nearly 4% by market close on release day [6]. Intelligence analyst Robert Lea, cited by Bloomberg, called the company's finances unsustainable, warning that "rising agentic AI will drive Z.ai's inference costs and losses higher" [6]. That skepticism sits next to Zhipu's own explanation for the delay - roughly two weeks for safety evaluation and hardening given the model's newly discovered exploit-chaining ability [3][4]. It's also not the first time an open Chinese coding model has raised this alarm: GLM-5.2 was already reported to be giving hackers a powerful new tool months before GLM-5.3 shipped [7], suggesting the pattern of open-weight capability outrunning safety review predates this specific release.

Historical Context

2026-02-12
Launched and open-sourced GLM-5, a 744B-parameter (40B activated) MoE model, expanding from GLM-4's 355B parameter scale.
2026-04
Released GLM-5.1 as part of its quarterly release cadence.
2026-06-25
The open-source GLM-5.2 model was already reported to be giving hackers a powerful new tool, foreshadowing the security concerns raised again with GLM-5.3.

Power Map

Key Players
Subject

Zhipu AI releases GLM-5.3 coding and cybersecurity model

ZH

Zhipu AI / Z.ai

Developer and publisher of GLM-5.3; a Hong Kong-listed Chinese AI lab positioning the model as the strongest open-weights coding system while delaying weight release for security review, shaping the narrative around open-weight cybersecurity risk from Chinese labs.

AN

Anthropic (referenced as Mythos 5)

Comparator model narrowly behind GLM-5.3 on CyberGym (83.8% vs 84.5%) but well ahead of it on ExploitBench (roughly high-70s% vs GLM-5.3's 54.4%), serving as the primary Western benchmark rival in cybersecurity coverage.

OP

OpenAI (GPT-5.6 Sol)

Second comparator referenced across CyberGym and ExploitBench scoring, used to benchmark GLM-5.3's relative cybersecurity strength against Western frontier labs.

CH

Chinese security teams / public vulnerability registry

Collaborators who worked with Zhipu to validate and disclose the 2,436 vulnerabilities the model surfaced across 269 open-source projects.

Fact Check

7 cited
  1. [1] Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
  2. [2] GLM-5.3 | Z.ai Documentation
  3. [3] Zhipu releases GLM-5.3 through its coding service, with weights still two weeks away
  4. [4] Zhipu AI Releases GLM-5.3, Claims It's the Strongest Open-Weights Coding Model
  5. [5] Zhipu Launches Flagship Model GLM-5.3 as China Seeks Mythos-Level Edge in Cyber Defence
  6. [6] GLM-5.3 Post-Training Produced Exploit Chains Z.ai Never Planned, Finds 1,097 Critical Bugs
  7. [7] China's Open-Source GLM-5.2 Is Already Giving Hackers a Powerful New Tool

Source Articles

Top 5

THE SIGNAL.

Analysts

Expressed skepticism about Z.ai's commercial sustainability despite GLM-5.3's technical claims, warning that rising inference costs from agentic AI workloads will keep pressuring the company's finances.

Robert Lea
Intelligence analyst (cited by Bloomberg)
The Crowd

Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: https://z.ai/blog/glm-5.3

@@Zai_org17850

GLM-5.3 isn't Fable or Sol level. But it's 743B and it lands 0.1 off Sol on Agents' Last Exam. I had a number in my head for where the open ones would stall out. I've moved it three times this year. The part I didn't expect is that they're pitching it as a security model.

@@ziwenxu_46

just launched GLM-5.3. And the most interesting part isn't that it's another 743B model. It's basically the same base model. They just kept pushing post-training. >Terminal-Bench 3.0: 4.6 -> 28.3 >DeepSWE: 46.2 -> 66.9 >ExploitBench: 24.4 -> 54.4

@@fanofaliens13

GLM 5.3 Released

@u/jmorant5551600
Broadcast
GLM 5.3 Is INSANE! The BEST Open Source Model EVER? BEATS MYTHOS? (Fully Tested)

GLM 5.3 Is INSANE! The BEST Open Source Model EVER? BEATS MYTHOS? (Fully Tested)

GLM-5.3 (Fully Tested): I GOT EARLY ACCESS & IT'S #1 ON MY BENCH!

GLM-5.3 (Fully Tested): I GOT EARLY ACCESS & IT'S #1 ON MY BENCH!

GLM-5.3 Released : Everything You Need to Know About Z.ai's New Coding Model (Benchmarks, Price )

GLM-5.3 Released : Everything You Need to Know About Z.ai's New Coding Model (Benchmarks, Price )