<p>This research presents a computationally efficient multimodal deep learning approach for accurately interpreting free-form natural language commands in autonomous vehicles (AVs). The system processes two primary inputs: the user’s voice commands and visual data captured by the vehicle’s camera. Its objective is to accurately identify referred objects in the voice command to determine the AV’s intended action or destination. We propose a novel architectural design that integrates a new early fusion technique, enhancing performance while maintaining low computational overhead. Our method is evaluated on the “Talk to Car” challenge dataset and outperforms existing approaches, offering a more accurate and cost-effective solution for real-world AV applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Command-based Object Detection with Early Fusion for Car (CODEF4Car)

  • Alireza Esna Ashari,
  • Vi Ly

摘要

This research presents a computationally efficient multimodal deep learning approach for accurately interpreting free-form natural language commands in autonomous vehicles (AVs). The system processes two primary inputs: the user’s voice commands and visual data captured by the vehicle’s camera. Its objective is to accurately identify referred objects in the voice command to determine the AV’s intended action or destination. We propose a novel architectural design that integrates a new early fusion technique, enhancing performance while maintaining low computational overhead. Our method is evaluated on the “Talk to Car” challenge dataset and outperforms existing approaches, offering a more accurate and cost-effective solution for real-world AV applications.