错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-stage Predictive World Model for Autonomous Driving

  • Xuwen Qiu,
  • Jinsong Liu,
  • Yikai Zhang,
  • Juan Deng

摘要

A predictive world model, which generates high-level representations of reality to forecast future states from historical and current observations, is crucial for autonomous driving. Most world models employ a single-depth decoder to predict all future states, overlooking the intrinsic differences between short-term and long-term forecasting. Aiming at this problem, we propose a Multi-Stage Decoding Framework based on the official ViDAR baseline. Our framework implements a hierarchical prediction strategy, assigning different temporal forecasting tasks to decoder layers at varying depths. Specifically, shallower layers are tasked with predicting the near future, while deeper layers handle the more challenging far-future prediction. This task specialization prevents the model from being overly complex for near-term tasks or too simplistic for long-term dependencies. Experiments with our multi-stage decoder on the OpenScenes-mini dataset achieve significant improvements, reducing the average Chamfer Distance (CD) by 17.2% for near-future and 40.3% for far-future predictions, demonstrating the effectiveness of our approach.