Infrared Small Target (IST) detection is characterized by the small size of the target and the lack of obvious features on the target surface. To increase the focus on target location information and reduce the impact of losing target information deep in the network on model performance. This paper proposes a densely nested infrared small target detection network based on coordinate joint channel attention. The network comprises three components: a feature extraction module utilizing dense nested architecture, a fusion module based on feature pyramid, and a detection module. Specifically, inspired by the DNANet and Coordinate Attention. The feature extraction section utilizes Coordinate Attention's capability to focus on target location information to compensate for DNANet's deficiency in this area. Simultaneously, channel attention is combined with Coordinate Attention to enhance the model's ability to focus on significant channel and location information. In addition, we utilize weighted Soft Intersection over Union (Soft-IoU) and Binary Cross-Entropy (BCE) loss to train the network to account for the model's ability to classify and locate targets. Finally, our method achieves better detection performance on the NUDT-SIRST dataset compared with advanced methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Densely Nested Infrared Small Target Detection Network Based on Coordinate Joint Channel Attention

  • Nengshuang Zhang,
  • Jing Zhang,
  • Huinan Guo,
  • Wuxia Zhang

摘要

Infrared Small Target (IST) detection is characterized by the small size of the target and the lack of obvious features on the target surface. To increase the focus on target location information and reduce the impact of losing target information deep in the network on model performance. This paper proposes a densely nested infrared small target detection network based on coordinate joint channel attention. The network comprises three components: a feature extraction module utilizing dense nested architecture, a fusion module based on feature pyramid, and a detection module. Specifically, inspired by the DNANet and Coordinate Attention. The feature extraction section utilizes Coordinate Attention's capability to focus on target location information to compensate for DNANet's deficiency in this area. Simultaneously, channel attention is combined with Coordinate Attention to enhance the model's ability to focus on significant channel and location information. In addition, we utilize weighted Soft Intersection over Union (Soft-IoU) and Binary Cross-Entropy (BCE) loss to train the network to account for the model's ability to classify and locate targets. Finally, our method achieves better detection performance on the NUDT-SIRST dataset compared with advanced methods.