<p>Car detection and ranging are critical components of autonomous driving systems, directly influencing driving safety. However, achieving optimal performance in both detection and ranging often requires substantial high-performance computing (HPC) resources, which can be computationally expensive. To address these challenges, we propose an efficient and streamlined visual assistance system that integrates an enhanced YOLO model with a transformer-based LiRT module and a simplified FSGBMN network for ranging. Our YOLO-LiRT model effectively combines low-level edge information with contextual data to improve detection performance, particularly in complex environments. Experimental results on the BDD100K and SODA10M datasets demonstrate impressive average precision scores of 44.8% and 63.9%, respectively. Furthermore, the FSGBMN network, utilizing the transformer-based LiRT module to capture contextual information, significantly reduces SGBM errors. Consequently, it achieves an average ranging MSE of 0.072 and 0.672, while maintaining real-time performance at 15 FPS on the Nvidia Jetson platform. A vision system for object detection and ranging requires high-performance computing (HPC) resources; our system can reduce the need for computing resources.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced car detection and ranging in adverse conditions using improved YOLO-LiRT and FSGBMN

  • Tao Luo,
  • Zhiwei Guan,
  • Dangfeng Pang,
  • Ruzhen Dou

摘要

Car detection and ranging are critical components of autonomous driving systems, directly influencing driving safety. However, achieving optimal performance in both detection and ranging often requires substantial high-performance computing (HPC) resources, which can be computationally expensive. To address these challenges, we propose an efficient and streamlined visual assistance system that integrates an enhanced YOLO model with a transformer-based LiRT module and a simplified FSGBMN network for ranging. Our YOLO-LiRT model effectively combines low-level edge information with contextual data to improve detection performance, particularly in complex environments. Experimental results on the BDD100K and SODA10M datasets demonstrate impressive average precision scores of 44.8% and 63.9%, respectively. Furthermore, the FSGBMN network, utilizing the transformer-based LiRT module to capture contextual information, significantly reduces SGBM errors. Consequently, it achieves an average ranging MSE of 0.072 and 0.672, while maintaining real-time performance at 15 FPS on the Nvidia Jetson platform. A vision system for object detection and ranging requires high-performance computing (HPC) resources; our system can reduce the need for computing resources.