World Labs introduced Atlas, a world model that understands text, images, video, and 3D information within one spatial context. Atlas can create scenes, reconstruct real spaces from only a few photographs, and simulate change over time. Its core strengths include precise camera control, 3D reconstruction, and real-to-sim applications for robotics. 🌍
1. Atlas and the Goal of a World Model
The world model pursued by World Labs is an AI capable of generating, reconstructing, and simulating any possible world. Its goal extends beyond producing one plausible image to understanding how a world looks, how it moves, and how it changes over time.
These capabilities can render imagined scenes for creators, virtualize real environments precisely, and help robots understand their surroundings and plan actions. World Labs describes this as foundational technology for spatial intelligence.
Atlas is a next-generation world model, an omni model pretrained at scale from the ground up. Instead of handling text, images, video, and 3D separately, it understands them together. It combines inputs into a shared spatial context, then generates the next image or scene. It maintains three-dimensional consistency with what it has already seen while naturally imagining and extending unseen areas.
World Labs says performance improved as Atlas scaled to more training compute and data, and expects this scaling effect to continue.
Atlas has four major application areas:
- Camera-controlled generation: Starting with one or more images, create images and video corresponding to desired camera positions and movements.
- Spatial reconstruction: Reconstruct real scenes from one to dozens of photographs and produce new views and explicit 3D outputs.
- Spatiotemporal simulation: Understand both spatial structure and temporal change in video for reconstruction and robot simulation.
- Image generation: Create ordinary images and 360-degree panoramas from text, following complex instructions and diverse visual styles.
Atlas will be incorporated into Marble, a World Labs product, and other future products.
2. Precise Camera Control and Scene Generation
Given one or more reference images, Atlas can generate a new scene from any user-specified camera position and angle. It does not simply make a similar picture: it preserves the content and geometry of the input as closely as possible while naturally imagining areas absent from the original.
After seeing a robot, for example, Atlas can create its unseen back. From a pool photograph, it may infer and extend a surrounding lawn outside the frame. This works across realistic scenes, visual styles, and camera movements.
Pixel-Level Camera Control
A major distinction is that Atlas does not rely only on vague natural-language directions for camera motion. It accepts precise camera geometry directly. Instead of a rough instruction such as "move the camera sideways," it composes a scene from camera data with a particular position, orientation, and path.
World Labs emphasizes that this gives users fine control over video composition and movement.
"You are directing a scene, not pulling the lever on a slot machine."
The workflow aims to let a creator plan and control the scene and camera like a director instead of leaving the result to chance.
Connecting Different Scenes Through Spatial Context
Like an LLM, Atlas first encodes its input as context and conditions generation on it. Atlas's context is not simply a sequence of sentences or images, however. It is a spatial context that includes where each image sits in 3D space.
Two unrelated images can therefore be placed inside a three-dimensional space and Atlas can generate one world connecting them. It imagines transitional elements such as doors, corridors, and narrow passages to join otherwise dissimilar locations.
One example uses a station corridor and a banquet hall as anchors, with Atlas generating the space between them.


Controllable Video up to One Minute Long
By combining camera movement with spatial-context management, Atlas can create longer videos. World Labs reports generating 1440p video up to one minute long from a small number of references and a human-designed camera path.
Atlas must preserve a consistent world throughout the movement. Rather than stitching together brief scenes, it constructs a continuous space through which the camera can travel.
3. Reconstructing Real Spaces from Few Images
Atlas reconstructs real objects and places from one or more input images. Conventional 3D reconstruction often requires special equipment or a dense set of many photographs; Atlas aims for faithful reconstruction from a few ordinary pictures.
World Labs describes this as meaningful progress on novel-view synthesis from sparse input images, a longstanding fundamental problem in 3D computer vision.
Imagine with Little Input, Become Precise with More
When part of a scene is absent from the photographs, Atlas fills the gap plausibly using learned knowledge of the world. When a particular real place must be reproduced accurately, users can supply more pictures.
The principle is simple: the more Atlas sees, the less it imagines. Two or three images can often produce a faithful reconstruction, while more than one hundred photographs can fit into the spatial context for greater precision.
"More inputs give Atlas more context. The more it sees, the less it imagines."
In one example, a single ground-level garden photograph lets Atlas generate an aerial view. The garden matches the input, while surrounding scenery is imagined. Adding a photograph of an outbuilding places both garden and outbuilding correctly, although the house on the left remains an estimate. Adding a photograph of the main house makes the full scene accurate.
A second example progressively reconstructs Stanford University's Main Quad. Between two and twenty-five ground-level photographs cover the lawn entrance through the colorful mosaics outside Memorial Church. With only those images, Atlas generates an aerial flight path over the campus.
Exploring the Same Place Along Different Paths
After reconstructing a space, Atlas can generate many camera paths through the same scene. Wherever the camera travels, the environment must remain the same place, while different speeds, lengths, and movement complexity create different moods.
One path may emphasize a particular object; another may move slowly to reveal the breadth of the space. A fixed 3D environment inferred from few inputs can therefore support diverse direction.
Point Clouds and 3D Gaussian Splats
Images and video are enough for some work, but robotics, games, design, and visual effects require explicit, manipulable 3D output. Because Atlas handles both image frames and depth maps, it can export a scene as a point cloud or 3D Gaussian splat.
A point cloud approximates geometry as a collection of spatial points. A 3D Gaussian splat adds color and shape information to represent a scene that can be rendered rapidly.
From one image, Atlas generates new views and estimates depth for each, constructing a 3D world. Given video of a real space, it predicts depth for every frame and combines them into a reconstruction. The model naturally fills regions never observed by the camera.
Completed Gaussian-splat scenes can render on-device at high resolution and frame rate. Marble uses the same representation, connecting Atlas naturally to World Labs' existing product system.
4. Video and Robot Simulation That Understand Time
Atlas is presented as a world simulator that handles not only spatial structure but how the world changes over time. This combination opens new uses in visual effects, robot learning, and simulation.
Bullet Time from Only a Few Cameras
Atlas can reconstruct footage from a few ordinary cameras as if it were captured in a professional studio with a large camera array. With footage from as few as three cameras, it can appear to freeze a moment and show it from an angle unavailable in the original recording.
"Real-world video can be reframed from new camera angles without an expensive capture studio."
The results did not rely on professional camera operators or specialized equipment. Researchers and engineers used tripods, clamps that fit in a bag, ordinary phones, and action cameras. Atlas reconstructs the scene from three to five camera views and lets users reposition the camera to recompose the footage.
Real-to-Sim That Generates a Robot's View
In robotics, spatial reconstruction is only half the task. As a virtual robot moves through a space, the simulation must generate what its cameras and sensors would observe. Along a path, Atlas produces both RGB images and depth data from the perspective of a body-mounted camera.
The virtual environment and the robot's view of it therefore come from the same Atlas model, rather than separate systems.
World Labs used phone video of two large environments and reconstructed each from only twenty-four frames. It then simulated different robot types along different paths and used Atlas to generate images from each robot's mounted camera.
Manipulation tasks involving robotic arms go further. A few casually captured videos help create simulations containing objects that move and interact. The resulting tasks can be varied easily by changing:
- Object types and positions
- Robot movement and actions
- Lighting and background
- Interactions among rigid, articulated, and deformable objects
This can produce diverse training data for robot learning and validation without repeatedly collecting large real-world datasets. It is the real-to-sim workflow emphasized by World Labs: constructing virtual environments from recordings of reality. 🤖
5. Image and 360-Degree Panorama Generation
Atlas's primary objective remains world modeling. World Labs treats every image as "a window into one possible world." Atlas nevertheless supports ordinary image generation, following complex text instructions, rendering text, and producing varied visual styles.
It also generates 360-degree images from text or image prompts, allowing creation of scenes that can be viewed in every direction rather than only flat images.
One example prompt is:
"At sunset, on a San Francisco beach near the Golden Gate Bridge. A rough chalk drawing with bold, heavy lines."
The example demonstrates Atlas following multiple visual requirements at once: location, time, medium, and line texture.
6. Unified Architecture and Performance Evaluation
An Architecture Centered on Spatial Context
Atlas handles several kinds of input and output in one unified architecture. World Labs says it designed a new foundation rather than reusing a conventional LLM or video architecture, placing spatial control at the center of the model.
Atlas is a multimodal autoregressive diffusion transformer:
- Multimodal: It directly handles text, images, camera position and orientation, and depth maps. Video is represented as a sequence of image frames. Each image and depth map is associated with an explicit camera pose, embedding spatial control in the architecture.
- Autoregressive: It processes an ordered sequence of data elements, referring to previous inputs and outputs as it creates the next element. Different input-output orderings let it adapt flexibly to many tasks.
- Diffusion: Beginning from a noise-like state, it removes noise over multiple steps to create an image or video. The number of inference steps can trade generation speed against quality.
- Transformer: Large matrix operations make the architecture well suited to modern AI hardware and provide a robust foundation for world modeling.


Inputs may include a text prompt, one or more photographs, and the pose of the camera that captured each image. Atlas integrates them as spatial context, then generates new image frames, complete video, and the depth information behind each scene.
"Atlas inputs are placed in 3D space to form spatial context, and the model generates multimodal outputs conditioned on that context."
Atlas combines strengths of LLMs with modern image and video models. As an autoregressive transformer it can use LLM optimizations such as KV caching, cache-aware routing, and disaggregated serving. As a latent diffusion model it can adopt advances including diffusion distillation, classifier-free guidance, noise-schedule adjustment, and improved VAE design.
Camera-Control and 3D-Reconstruction Benchmarks
Because Atlas is general-purpose, no single score can express all its capabilities. World Labs quantitatively evaluated two tasks: camera-conditioned generation and 3D reconstruction from sparse inputs.
For camera-conditioned generation, the model received one input image and one to three cinematic camera movements, such as pans, truck movements, and crane movements. Atlas received the path in its native precise camera-input format; comparison video models could not process it directly and instead received text descriptions of the movement.
External human raters judged which model followed the intended path more closely. World Labs reports that Atlas outperformed recent video models and that its advantage grew as camera paths became more complex.
Atlas was preferred at the following rates:
- Versus MiniMax H3: 75 percent
- Versus Gemini Omni Flash: 81 percent
- Versus Happy Horse 1.1: 86 percent
- Versus FLUX 3: 93 percent
- Versus Seedance 2.5: 94 percent
For 3D reconstruction, the model received multiple images and camera poses and predicted a 3D point corresponding to every input pixel. World Labs says it reproduced all baseline outputs under the same evaluation procedure for a fair comparison.
It claims that despite being a general model for both generation and reconstruction, Atlas outperformed leading specialized open-source 3D reconstruction models. Lower reconstruction error is better in this evaluation.
Capabilities That Grow with Scale
World Labs attributes much of modern AI progress to scaling data and compute. Atlas was pretrained from scratch on large, diverse multimodal data, and the team trained models at multiple sizes and compute levels during development.
It says each increase in compute tended to produce new capabilities and expresses confidence that continued scaling will substantially improve future world models.
7. Early Access and Future Plans
Atlas is now in early access with selected partners. Developers and companies seeking to build products or services with Atlas can request access.
World Labs aims to make Atlas the definitive world model for generating, reconstructing, and simulating any world. It is also hiring researchers and engineers to advance spatial intelligence.
Closing
Atlas goes beyond image or video generation toward treating space, cameras, depth, and time as one consistent world. Its defining capabilities are reconstructing 3D environments from sparse real-world data, directing scenes through desired cameras, and generating the virtual worlds a robot would observe. How far performance continues to expand through scaling will be a key test of World Labs' vision for spatial intelligence.
