A Comprehensive Review of Vision Language Models in the Realm of Full Self-Driving Systems
摘要
This paper presents a comprehensive review of the advancements in the context of autonomous driving (AD) systems, exploring the potential and current applications of Large Language Models and Vision Language Models in enhancing Full Self-Driving (FSD) capabilities. As AD systems evolve, the integration of multi-modal inputs, particularly visual and textual data, through advanced VLMs has become a critical factor in improving vehicle perception, decision-making, and interaction with complex environments. This survey addresses six key research questions that assess the progress of VLMs in various aspects, including object detection, scene understanding, natural language instructions, and safety-critical decision processes in FSD systems. By synthesizing recent developments and ongoing research, the objective of this paper is to provide a clear overview of the current state of VLM integration in autonomous driving and to outline future directions for advancing these technologies in real-world applications.