The Architectural Bet: NEO-unify

SenseTime's core wager with U1 Pro is architectural, not just a bigger model. NEO-unify, first introduced through a March 2026 research collaboration with NTU [1], unifies language and visual representation inside a single 'inner kernel' rather than bolting a separate image-generation module onto a text-focused LLM [3]. Per SenseTime's own architecture description on the model's GitHub repository, NEO-unify 'eliminates both Visual Encoder (VE) and Variational Auto-Encoder (VAE)' by design - a first-principles choice rather than an incremental fix [4]. That same architecture underlies both the open-source base U1 models released in April 2026 [2]and the newer flagship U1 Pro announced in July [3], though the VE/VAE-elimination claim specifically describes the NEO-unify design rather than being independently confirmed for each downstream variant. The bet appears to be paying off in adoption terms for the base model at least: daily image-generation volume on U1 roughly tripled between May and June 2026, and combined GitHub stars across U1 and SenseNova-Skills passed 8,000 by July [8].



