Enhancing Interpretable World Models for End-to-End Driving via Functional Decomposition
摘要
The autonomous driving community has seen significant advancements in end-to-end frameworks that map raw sensor input directly to vehicle motion plans, bypassing traditional modularized pipelines. While joint optimization improves efficiency, the deployment of these systems in safety-critical environments remains constrained by challenges in generalization, interpretability, and controllability. In this work, we propose enhancing end-to-end world models through structured decomposition into interpretable functional decomposition—each associated with distinct perception, reasoning, and decision-making elements. By systematically mapping functional decomposition to scenario constraints and interpretability metrics, our approach enables modular understanding, enforces safety requirements, and supports automated scene validation. This framework improves transparency, operational safety, and scalability without compromising the advantages of data-driven learning.