Advancing End-to-End Autonomous Driving with Visual Guidance and Simulated Interaction
摘要
Environmental complexity results in suboptimal performance and interpretability issues when using direct mapping via imitation learning. This paper proposes an end-to-end model based on imitation learning within a multi-task framework. The model enhances environmental understanding and interpretability from two aspects: visual guidance and simulated environmental interaction. For visual guidance, semantic segmentation is used as a visual prior to learn the intuitive state of the environment. This approach helps the model understand spatial layout and semantic content, providing a better basis for decision-making. For simulated environmental interaction, a self-regressive network utilizing Gated Recurrent Units(GRU) and Kolmogorov-Arnold Networks(KAN) is introduced. By integrating information from different branches, temporal modules are built to simulate environmental interactions, enhancing the robustness and interpretability of the model’s outputs. The effectiveness of method is evaluated through various scenarios in the CARLA simulator, demonstrating its superiority in handling complex driving conditions.