<p>To address the issue of low accuracy and slow detection speed in multi-object detection and distance estimation for autonomous driving systems operating in complex environments, this study proposes an enhanced multi-object detection algorithm based on YOLOv8 framework. First, an early cross-modal fusion method based on Vision Transformer was implemented to integrate visual, millimeter-wave radar, and LiDAR data for enhanced feature extraction. Subsequently, we replaced the original CSPDarknet backbone in YOLOv8 with a lightweight CSPHet backbone. Concurrently, a dual-attention mechanism combining HS-FPN and CCFM was developed to improve small-object feature extraction. Furthermore, the improved YOLOv8 model was integrated with the DeepSORT algorithm to achieve robust multi-object tracking. Additionally, A monocular ranging model also was employed to detect and estimate distances to forward-facing pedestrians and vehicles. Experimental results on a benchmark dataset demonstrate superior performance: the mAP for vehicle and pedestrian exceed those of the BEVFormer method by + 3.3% and + 1.1%, respectively, and surpass the standard YOLOv8 by 5.2% in mAP50, achieving 86.8% recall and 77.5% mAP50. These results show that the proposed algorithm accurately detects, tracks, and estimates the distance of pedestrians and vehicles in real-time driving scenarios, enabling reliable autonomous decisions and robust early warnings in complex and uncertain environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CSPHet-CCFM-YOLOv8: A Lightweight Multi-object Detection and Distance Estimation Framework for Autonomous Driving Systems

  • Tongying Guo,
  • Siwen Yu

摘要

To address the issue of low accuracy and slow detection speed in multi-object detection and distance estimation for autonomous driving systems operating in complex environments, this study proposes an enhanced multi-object detection algorithm based on YOLOv8 framework. First, an early cross-modal fusion method based on Vision Transformer was implemented to integrate visual, millimeter-wave radar, and LiDAR data for enhanced feature extraction. Subsequently, we replaced the original CSPDarknet backbone in YOLOv8 with a lightweight CSPHet backbone. Concurrently, a dual-attention mechanism combining HS-FPN and CCFM was developed to improve small-object feature extraction. Furthermore, the improved YOLOv8 model was integrated with the DeepSORT algorithm to achieve robust multi-object tracking. Additionally, A monocular ranging model also was employed to detect and estimate distances to forward-facing pedestrians and vehicles. Experimental results on a benchmark dataset demonstrate superior performance: the mAP for vehicle and pedestrian exceed those of the BEVFormer method by + 3.3% and + 1.1%, respectively, and surpass the standard YOLOv8 by 5.2% in mAP50, achieving 86.8% recall and 77.5% mAP50. These results show that the proposed algorithm accurately detects, tracks, and estimates the distance of pedestrians and vehicles in real-time driving scenarios, enabling reliable autonomous decisions and robust early warnings in complex and uncertain environments.