To address edge blurring and semantic noise in small target detection within complex remote sensing scenarios, this study proposes FSDA-YOLO, a frequency-spatial dynamic perception framework. Based on YOLOv11s, the method constructs a frequency-aware transformer network (FAT-Net) via Fourier transform to decouple high-frequency edge features from low-frequency backgrounds, enhanced by spectral multi-head attention (S-MHA) for critical frequency band focusing. A multi-granularity feature pyramid network (MGFPN) integrates a 160 × 160 ultra-high-resolution branch (UHR-Branch) with bidirectional cross-layer interactions to strengthen multi-scale representation. A focal-weighted loss function optimizes hard sample training. Experimental results show 2.91 × 105 parameter reduction with 2.0% AP and 5.6% AP50 improvements on VisDrone2019, alongside 3.4% AP50 gain on Tiny Person, demonstrating robust sub-10-pixel target detection capability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FSDA-YOLO: Frequency-Aware Multi-scale Fusion for Small Object Detection

  • Haoyu Guo,
  • Yude Wang,
  • Dezhong Jing,
  • Xinyu Wang,
  • Teng Liu,
  • Fei Song,
  • Yuanpei Wang

摘要

To address edge blurring and semantic noise in small target detection within complex remote sensing scenarios, this study proposes FSDA-YOLO, a frequency-spatial dynamic perception framework. Based on YOLOv11s, the method constructs a frequency-aware transformer network (FAT-Net) via Fourier transform to decouple high-frequency edge features from low-frequency backgrounds, enhanced by spectral multi-head attention (S-MHA) for critical frequency band focusing. A multi-granularity feature pyramid network (MGFPN) integrates a 160 × 160 ultra-high-resolution branch (UHR-Branch) with bidirectional cross-layer interactions to strengthen multi-scale representation. A focal-weighted loss function optimizes hard sample training. Experimental results show 2.91 × 105 parameter reduction with 2.0% AP and 5.6% AP50 improvements on VisDrone2019, alongside 3.4% AP50 gain on Tiny Person, demonstrating robust sub-10-pixel target detection capability.