The real world is always in motion. To operate autonomously, physical AI systems — including robots, autonomous vehicles (AVs) and smart spaces — need to understand not just what they see and what caused that to happen, but what’s likely to happen next. In a warehouse, a robot may encounter object configurations it’s never seen before. On the road, an AV may need to respond when a pedestrian steps out from between parked cars. And in a factory, a safety system must predict where a forklift is heading, not just detect that it’s there. Capturing and recreating those scenarios in the real world is slow, expensive and often impossible to repeat at scale. NVIDIA Cosmos 3 is built for that loop. The new world foundation model — announced today at NVIDIA GTC Taipei at COMPUTEX — combines vision reasoning and multimodal generation across text, video, images, ambient sound and action in a single model to help developers create world data with physical context. Cosmos 3 powers perception, prediction and action. Learn more about how Cosmos 3’s mixture-of-transformers architecture enables a reasoning block to first interpret what is happening in a scene, then harnesses a generation block to use that context to create physically grounded outputs, from synthetic video to robot-task data. Cosmos 3 is a generalist foundation model trained on diverse data that gives it a broad understanding of how scenes, motion and robotic actions relate. It’s an omnimodel with native action generation, meaning it can produce numerical action data, such as joint angles, gripper positions and trajectory points, that describe how a robot should move to complete a task. In order to learn, robots need more than images or videos of a scene. For pick-and-place tasks, for example, they need action signals that guide how to reach, grasp, move