NVIDIA released Cosmos 3, an open-source world foundation model for physical AI built on a mixture-of-transformers architecture that combines vision reasoning, world generation, and action prediction in a single system. The model is available in three sizes under the Linux Foundation's OpenMDW-1.1 license: Cosmos 3 Super (64B parameters) for high-fidelity world modeling, Cosmos 3 Nano (16B) for efficient reasoning and deployment, and Cosmos 3 Edge (4B, announced but not yet released) for on-device robotics and edge inference. The model was trained on 20 trillion tokens of multimodal data, including nearly 1 billion images, 400 million videos, ambient audio, and robot action trajectories.
Cosmos 3 ranks #1 on Artificial Analysis for open-weights text-to-image and image-to-video generation, leads PAI-Bench for world generation and Physics-IQ for image-to-video, and ranks #1 on RoboLab for robot policy. Cosmos 3 Super is also the highest-ranked open model on VANTAGE-Bench for vision understanding across fixed-camera warehouse, transportation, and smart-space footage. Unlike prior video-only foundation models, Cosmos 3 natively outputs robot action signals—joint angles, gripper positions, trajectory points—enabling direct policy training and synthetic data generation for embodied AI systems.
NVIDIA launched the NVIDIA Cosmos Coalition, a global collaboration with robotics and AI leaders including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI to contribute models, research, and evaluation methods. Developers can post-train Cosmos 3 on proprietary data and hardware, and deployment partners including Baseten, CoreWeave, Microsoft Azure, and Deep Infra offer containerized inference. For architects, Cosmos 3 offers a single unified model replacing separate pipelines for world reasoning, simulation, and policy training—reducing development cycles from months to days.