The efficiency pitch, and its hidden caveats
Reflection AI's headline claim is that Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using three to four times less inference compute [1]. But Reflection itself describes that figure as an approximate comparison rather than a measured cost [1], and outside technical analysis notes the comparison appears to exclude prompt prefill, context-dependent attention, and serving overhead [3]- the kind of real-world costs that determine what a model actually costs to run in production, not just in a benchmark table. The efficiency argument rests on Beam's mixture-of-experts design: only 23 billion of its 501 billion parameters activate per token, versus 49 billion active parameters in DeepSeek's V4 Pro [1]. That smaller active footprint is real and verifiable from the architecture itself; the 3-4x compute-savings framing against GLM-5.2 is not, at least not yet, since independent testers have only had vendor-granted access ahead of the public weight release [3].


