Guangzhou, April 29, 2026 — XPENG (NYSE: XPEV, HKEX: 9868), a leading China-based high-tech company, recently officially released its X-World Technical Report, providing a comprehensive breakdown of the model's construction and deployment across data, architecture, training, validation, and application. X-World is a controllable, multi-view generative world model designed for autonomous driving. Built on video diffusion technology, it features real-time response and continuous generation capabilities across multiple perspectives. The report highlights X-World's practical value within XPENG's autonomous driving ecosystem, where it is already integrated into production workflows such as closed-loop simulation, online reinforcement learning, and data synthesis. Furthermore, during the recent rollout of VLA 2.0 to users, X-World has been extensively utilized for environmental simulation and model evaluation throughout the R&D and validation phases. The evaluation of autonomous driving systems primarily relies on real-world road testing and simulation testing. Among these, simulation testing possesses advantages such as lower costs, higher efficiency, broader scenario coverage, and repeatable verification. Traditional simulation evaluation extensively adopts technical roadmaps based on 3D Gaussian Splatting (3DGS). While these methods can reproduce real-world scenes to a certain extent, they often struggle to effectively generate and evaluate subsequent scenes beyond the existing reconstruction range when an autonomous driving model produces behaviors that significantly deviate from the original collected trajectory, such as sharp lane changes or detours. Consequently, the industry still relies heavily on real-vehicle road testing, a method characterized by high costs, limited scenario coverage, and the difficulty of reproducing specific situations. To resolve these bottlenecks, the XPENG Generative World Model team sought to build a "real-world simulator" capable of generating future videos that comply with physical constraints under given action conditions, while maintaining high controllability and stability throughout the continuous generation process. In this context, X-World was born. By inputting multi-camera historical video streams and the driving actions