<p>Detecting small, dense, and occluded objects in UAV-based remote sensing imagery is a crucial challenge, requiring algorithms that balance high accuracy with real-time efficiency. To address this, we introduce SD-YOLO, a model enhancing YOLOv8 through three key innovations. First, its lightweight design prunes redundant low-resolution feature maps and adds a tiny detection head, which reduces parameters considerably. Second, the backbone is enhanced by replacing standard C2f blocks with our C2f-DMSC for superior multi-dimensional feature extraction, and by integrating a Transformer module to capture global context. Third, our MSCBAM attention module expands the receptive field and refines feature processing by emphasizing critical regions. To meet diverse application needs, we offer two variants: the highly efficient SD-YOLOn and the high-accuracy SD-YOLOs, created via channel scaling. Evaluations show SD-YOLOn achieves <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(35.8\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>35.8</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(76.3\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>76.3</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(43.7\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>43.7</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\texttt {mAP}_{0.5}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="monospace">mAP</mi> <mrow> <mn>0.5</mn> </mrow> </msub> </math></EquationSource> </InlineEquation> results on VisDrone-2019, LEVIR-Ship, and DOTA, respectively, with a model size three times smaller, thus demonstrating its effectiveness for small, dense object detection in remote sensing.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SD-YOLO: A lightweight and high-performance deep model for small and dense object detection

  • Phuc-Thinh Huynh,
  • Minh-Thanh Le,
  • Tran Duc Tan,
  • Thien Huynh-The

摘要

Detecting small, dense, and occluded objects in UAV-based remote sensing imagery is a crucial challenge, requiring algorithms that balance high accuracy with real-time efficiency. To address this, we introduce SD-YOLO, a model enhancing YOLOv8 through three key innovations. First, its lightweight design prunes redundant low-resolution feature maps and adds a tiny detection head, which reduces parameters considerably. Second, the backbone is enhanced by replacing standard C2f blocks with our C2f-DMSC for superior multi-dimensional feature extraction, and by integrating a Transformer module to capture global context. Third, our MSCBAM attention module expands the receptive field and refines feature processing by emphasizing critical regions. To meet diverse application needs, we offer two variants: the highly efficient SD-YOLOn and the high-accuracy SD-YOLOs, created via channel scaling. Evaluations show SD-YOLOn achieves \(35.8\%\) 35.8 % , \(76.3\%\) 76.3 % , and \(43.7\%\) 43.7 % \(\texttt {mAP}_{0.5}\) mAP 0.5 results on VisDrone-2019, LEVIR-Ship, and DOTA, respectively, with a model size three times smaller, thus demonstrating its effectiveness for small, dense object detection in remote sensing.