One Model Now Does What Used to Take a Pipeline
World Labs built Atlas as a single multimodal autoregressive diffusion transformer, pretrained from scratch to natively process text, images, video, and 3D within one spatial context, collapsing what used to require separate photogrammetry, splat-training, and camera-solving pipelines into one model [1]. The model generates up to a minute of 1440p video with pixel-perfect camera control and can output explicit 3D point clouds and Gaussian splats alongside standard video frames [1]. On sparse-view 3D reconstruction, World Labs reports Atlas hits a 25.3 mean absolute-relative pointmap error versus 28.7 for the next-best specialist model, and industry analysis has framed the result as evidence that generation and reconstruction are fundamentally the same task [2]. That framing is the real news here: Atlas is not pitched as a better video generator so much as a bet that one sufficiently large model can absorb an entire category of specialist 3D tools.



