August 27, 2026 – At its Physical AI Sharing and Second-Generation VLA Experience Day, XPeng officially announced the debut of its X-Foresight predictive world model in production vehicles. This advanced system can predict events up to six seconds in advance and anticipate the potential behaviors of surrounding traffic participants. The newly upgraded second-generation VLA model features comprehensive enhancements, boosting multi-dimensional comprehensive safety capabilities by 20 times.
The second-generation VLA introduces four core upgrades. Notably, the on-device model parameters have increased by 3.5 times, making it over 15 times larger than mainstream VLA models. This massive parameter scale delivers superior generalization capabilities, laying the foundation for rapid expansion into global markets. Additionally, end-to-end response speed has improved by 300%, enabling millisecond-level reactions and rapid evasive maneuvers. The system now supports up to 30 seconds of effective temporal sequence, marking the first time long-term temporal memory has been integrated into autonomous driving decision-making.

On real-world roads, unpredictable scenarios such as sudden braking, pedestrians changing direction, or unexpected lane cut-ins are common. To address these, XPeng’s X-Foresight model utilizes Flow Matching technology to predict multiple possible future outcomes, allowing the vehicle to select the optimal course of action. Rather than relying solely on past and current visual inputs, the model predicts based on the target’s position, speed, movement trends, and road environment. By generating multiple driving intentions through long-term temporal representations, it transitions from a single deterministic path to multiple possibilities, ultimately using reinforcement learning to create more natural and human-like trajectories.
Complementing this is the newly introduced Infini-VLA long-term temporal architecture, which enables the vehicle to “remember” the past 30 seconds of its environment, feeding more valid information into the decision-making model. XPeng emphasized that these technologies work in unison: Infini-VLA brings in the past, streaming autoregressive inference processes the present, and X-Foresight with Flow Matching anticipates the future. Together, they form a complete spatiotemporal decision-making loop.
From a technical evolution standpoint, the second-generation VLA was first unveiled at XPeng Tech Day on November 5, 2025, and officially rolled out on March 2, 2026. As a native multimodal physical world foundation model integrating vision, hearing, and reading, it innovatively eliminates the traditional “language translation” step, achieving direct end-to-end generation from visual signals to action commands. Trained on nearly 100 million video clips—equivalent to tens of thousands of years of human driving experience—the model requires no manual data annotation.
