The Three Walls: Why Physical AI's Data Problem Dwarfs Language Models'
Zhiyuan partner Yao Maoqing gave WAIC 2026's clearest framing of what stands between today's demos and commercially reliable robots: a data wall, a representation wall and a closed-loop wall, in his words the model determines the starting point but data determines the endgame [1]. The scale problem is stark. According to figures shared at the AGIBOT-hosted forum, physical AI's current data volume is only about 1/20000 of what large language models trained on, and reaching a genuine physical-world ChatGPT moment would require at least 100 million hours of real interaction data [2]. That is a fraction of the roughly 100 billion hours believed to underlie LLM training, but still an enormous gap from where the field sits today. Zhiyuan and Embodied Technology's response is to open-source infrastructure rather than hoard it - their AGIBOT WORLD dataset, described as the industry's first million-scale real robot dataset, has already logged more than 1.2 million cumulative downloads. Meanwhile Georgia Tech's Xu Danfei reported his team has independently accumulated 20,000 hours of human first-person-view data and observed something resembling a linear scaling law, evidence that more data reliably buys more capability even if nobody has enough of it yet.




