<p>Container damage detection is a critical process for ensuring cargo safety, improving logistics efficiency, and reducing repair costs. To address the issues of susceptibility to environmental interference and high false detection rates encountered when applying the YOLOv8 model to complex scenarios, we propose an improved YOLOv8 model named EB-YOLOv8 based on its original architecture. This model replaces the original backbone network with EfficientViT, a memory-efficient vision transformer with cascaded group attention, to enhance feature extraction capabilities. It introduces the Bidirectional Feature Pyramid Network (BiFPN) to improve multi-scale feature fusion via bidirectional connections and adaptive weighted fusion mechanisms. Additionally, it employs the Enhanced Intersection over Union (EIoU) and Varifocal Loss to optimize bounding box regression and classification confidence prediction, respectively. We conducted extensive ablation experiments and comparative experiments using a real-world port dataset. Experimental results demonstrate that EB-YOLOv8 achieves substantially better detection accuracy and positioning performance for various container damage types under complex backgrounds compared to the original YOLOv8. It obtains an mAP50 of 87.99%, which is 8.08% points higher than that of YOLOv8s, and reaches an FPS of 41.16, which meets the requirements of real-time detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An improved YOLOv8 model for container damage detection in complex scenarios

  • Yuxuan Zhang,
  • Zhenghao Tang,
  • Yinghan Yang,
  • Aiping Zhou

摘要

Container damage detection is a critical process for ensuring cargo safety, improving logistics efficiency, and reducing repair costs. To address the issues of susceptibility to environmental interference and high false detection rates encountered when applying the YOLOv8 model to complex scenarios, we propose an improved YOLOv8 model named EB-YOLOv8 based on its original architecture. This model replaces the original backbone network with EfficientViT, a memory-efficient vision transformer with cascaded group attention, to enhance feature extraction capabilities. It introduces the Bidirectional Feature Pyramid Network (BiFPN) to improve multi-scale feature fusion via bidirectional connections and adaptive weighted fusion mechanisms. Additionally, it employs the Enhanced Intersection over Union (EIoU) and Varifocal Loss to optimize bounding box regression and classification confidence prediction, respectively. We conducted extensive ablation experiments and comparative experiments using a real-world port dataset. Experimental results demonstrate that EB-YOLOv8 achieves substantially better detection accuracy and positioning performance for various container damage types under complex backgrounds compared to the original YOLOv8. It obtains an mAP50 of 87.99%, which is 8.08% points higher than that of YOLOv8s, and reaches an FPS of 41.16, which meets the requirements of real-time detection.