<p>Small object detection poses a notable challenge in object detection tasks. Previous small object detection methods process features on the entire image, while spatial dominance of background pixels leads to the progressive attenuation of small object features through convolutional layers, failing to generate discriminative representations, causing confusion between objects and background. To this end, we propose a novel semantic decoupling module, which explicitly separates the semantic information space into two distinct components: an object feature space and a background feature space. In their respective spaces, the object representation is enhanced while background features are simultaneously constrained through proper mapping strategy. By transfer learning, we design a multi-projector network with multi-views to disentangle irrelevant background and object correlations, ensuring consistent feature encoding across complex backgrounds and mitigating background confusion in downstream tasks. Extensive experiments conducted on the Pascal VOC, Traffic Light, and Rail Worker datasets demonstrate a marked reduction in background errors and a substantial improvement in detection performance compared to existing methods. Our code is available: https://github.com/gaohua-1/TL_OBSDM DOI:10.5281/zenodo.14613329.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic decoupling and transfer learning for enhanced small object detection

  • Gaohua Liu,
  • Jinghao Zhang,
  • Junhuan Li,
  • Shuxia Yan,
  • Xiangyu Kong,
  • Rui Liu,
  • Yueyang Li

摘要

Small object detection poses a notable challenge in object detection tasks. Previous small object detection methods process features on the entire image, while spatial dominance of background pixels leads to the progressive attenuation of small object features through convolutional layers, failing to generate discriminative representations, causing confusion between objects and background. To this end, we propose a novel semantic decoupling module, which explicitly separates the semantic information space into two distinct components: an object feature space and a background feature space. In their respective spaces, the object representation is enhanced while background features are simultaneously constrained through proper mapping strategy. By transfer learning, we design a multi-projector network with multi-views to disentangle irrelevant background and object correlations, ensuring consistent feature encoding across complex backgrounds and mitigating background confusion in downstream tasks. Extensive experiments conducted on the Pascal VOC, Traffic Light, and Rail Worker datasets demonstrate a marked reduction in background errors and a substantial improvement in detection performance compared to existing methods. Our code is available: https://github.com/gaohua-1/TL_OBSDM DOI:10.5281/zenodo.14613329.