The Forensic Trail: Fingerprinting a Model With No Name
Ox Alpha showed up on OpenRouter on August 20, 2026, listed only under the generic 'stealth' provider category, with no company name, logo, or press release attached [1]. That vacuum turned identity-guessing into a public research project. On August 22, a researcher going by Chetaslua triggered a Java stack trace from the model's API that exposed an internal class path matching a documented Zhipu API route, calling it implementation detail leaking at the routing layer rather than another round of community guessing [2][3]. Separate tokenizer testing found Ox Alpha's token counts run a consistent 75 tokens higher than Zhipu's GLM-5.3 across roughly 30 prompts spanning 14 writing systems, a fingerprint that is hard to fake by accident [3]. Testers pushed the comparison into more sensitive territory too: one live test asked Ox Alpha whether Taiwan is part of China and got an answer nearly identical to GLM's, another data point some read as circumstantial confirmation of the Zhipu link. On the Kingbench benchmark, Ox Alpha scored 87.5%, trailing GLM-5.3's 91.25% by less than four points and clearing Opus 4.8 and Qwen 3.8 Max by a wide margin [4]. Independent analysis of its behavior also estimated a roughly 744-billion-parameter mixture-of-experts architecture with about 40 billion active parameters, though this figure is inferred rather than disclosed [5]. Coding benchmarks told a messier story: an initial 10-question DeepSWE sample had Ox Alpha beating Fable 5, GLM-5.3, and GPT-5 outright at an 80% pass rate [5][10], but a later 113-task run from the same tester found a more modest 63% pass rate, described as 'on par with GPT-5.6 Sol mid' rather than clearly ahead of it [6]. Small benchmark samples make for viral claims and shakier conclusions. Not every side-by-side test lines up neatly behind Zhipu, either: one tester running personal coding-challenge benchmarks noted stylistic similarities to GLM's known quirks but ultimately ruled out GLM, Gemini, GPT, and DeepSeek as matches based on behavioral fingerprints, concluding the identity remains genuinely ambiguous even under close hands-on testing.


