Drone aerial imagery often features densely distributed targets with indistinct features and significant scale variations, leading to challenges such as missed detections and false positives in object detection. To address these issues, this study proposes an innovative lightweight small object detection framework tailored for high-density, scale-variant scenarios. The framework incorporates an efficient small object detection layer and a novel weighted multi-branch fusion neck architecture, which synergistically optimize feature extraction and multi-scale information fusion for high-performance small object detection. First, the head architecture refines the original P5 detection layer by replacing it with a P2 small object detection layer, reducing network depth and computational overhead while enhancing the efficiency and accuracy of detecting small objects in densely distributed aerial images. Second, the backbone architecture integrates Spatial-Depthwise Convolution (SPDConv), a structure specifically suited for low-resolution images, effectively preserving fine-grained details and discriminative feature information critical for identifying small targets. Finally, the neck architecture introduces a novel Weighted Multi-Branch Supportive FPN (WMSFPN), which fuses shallow spatial and deep semantic information through multiple additional branches. By incorporating learnable weight parameters, WMSFPN dynamically balances the contributions of different layers, enabling efficient multi-scale information integration and significantly enhancing detection performance for small objects. Experimental results on the VisDrone2019 dataset demonstrate that, compared to YOLOv8n, the proposed framework improves mAP50 from 33.4% to 40.0%, a 6.6% increase, while reducing the number of parameters from 3.0M to 1.0M (1M = 106), a 66.7% decrease. The model also outperforms related lightweight methods, achieving higher accuracy, real-time detection capability, and superior generalization across multiple small object datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Small Object Detection in Drone Imagery: A Lightweight Weighted Multi-branch Supportive Fusion

  • Mingfei Rong,
  • Zhen Wang,
  • Zunwei Fu

摘要

Drone aerial imagery often features densely distributed targets with indistinct features and significant scale variations, leading to challenges such as missed detections and false positives in object detection. To address these issues, this study proposes an innovative lightweight small object detection framework tailored for high-density, scale-variant scenarios. The framework incorporates an efficient small object detection layer and a novel weighted multi-branch fusion neck architecture, which synergistically optimize feature extraction and multi-scale information fusion for high-performance small object detection. First, the head architecture refines the original P5 detection layer by replacing it with a P2 small object detection layer, reducing network depth and computational overhead while enhancing the efficiency and accuracy of detecting small objects in densely distributed aerial images. Second, the backbone architecture integrates Spatial-Depthwise Convolution (SPDConv), a structure specifically suited for low-resolution images, effectively preserving fine-grained details and discriminative feature information critical for identifying small targets. Finally, the neck architecture introduces a novel Weighted Multi-Branch Supportive FPN (WMSFPN), which fuses shallow spatial and deep semantic information through multiple additional branches. By incorporating learnable weight parameters, WMSFPN dynamically balances the contributions of different layers, enabling efficient multi-scale information integration and significantly enhancing detection performance for small objects. Experimental results on the VisDrone2019 dataset demonstrate that, compared to YOLOv8n, the proposed framework improves mAP50 from 33.4% to 40.0%, a 6.6% increase, while reducing the number of parameters from 3.0M to 1.0M (1M = 106), a 66.7% decrease. The model also outperforms related lightweight methods, achieving higher accuracy, real-time detection capability, and superior generalization across multiple small object datasets.