ADAPT: Action-Aware Driving Caption Transformer
摘要
Benefiting from precise perception, real-time prediction and reliable planning, autonomous driving systems have exhibited exceptional performance in research. However, the high complexity and opacity prevent its application in practice. To introduce a user-friendly autonomous driving system, we propose a driving captioner to generate real time description and explanation of self-driving systems in natural language. Specifically, we unify the end-to-end autonomous driving and video captioning tasks into a single yet effective framework by introducing an additional captioning head to describe the action of the vehicle and explain the reasons. Besides, we exploit an effective accelerating method to accelerate the inference process, which decreases the average inference time from 0.670 s to 0.298 s. Through extensive experiments on both simulation datasets and real-world datasets, we show the superior generalization ability and robustness of the proposed framework.