Fei-Fei Li's world model takes the stage: Atlas generates frames with pixel-perfect camera control
World Labs introduced Atlas, a multimodal world model trained from scratch: it generates image and video frames with pixel-perfect camera control and reconstructs the scene in 3D.
World Labs, the company of Fei-Fei Li, one of the founding names of modern AI, has shown the big thing it has been building: Atlas, in the company's words a first-of-its-kind multimodal world model trained from scratch. Li announced it as a major milestone for the team.
Atlas's claim has two layers. On the generation side, the model produces image and video frames with pixel-perfect camera control: give it a single input image and a camera path, and it renders frames as if shot along that route. On the understanding side, it reconstructs what it generates in 3D, meaning the scene lives in the model's head as a walkable space, a flat picture no longer.
World models are widely seen as the next front of the agent race: after models that mastered language come models that master space, physics and object permanence. Arriving the same week as Google's AgentHands is no coincidence; one reaches a virtual hand into real space, the other carries space itself into the model. Games, film, robotics and simulation will be reshaped between those two ends.