<p>Fine-grained ship classification (FGSC) in remote sensing imagery is a critical yet challenging task due to significant scale variations, complex backgrounds, high inter-class similarity, and intra-class variability. While advanced deep learning methods have improved performance, a major limitation of many existing approaches is their dependence on costly auxiliary annotations like bounding boxes or part labels, which restricts their scalability and practical deployment. To overcome these issues, we propose a novel FGSC framework for remote sensing imagery that requires only image-level supervision. Our method leverages a powerful convolutional backbone (ConvNeXtV2) enhanced with an integrated cascaded attention mechanism that performs progressive channel-wise refinement and spatial highlighting to improve feature discriminability by dynamically focusing on salient information. Furthermore, it incorporates an innovative weighted bidirectional fusion strategy to effectively combine features from selected high-level network stages, thereby enhancing robustness to scale variations and complex scenes. Extensive experiments on challenging FGSC datasets demonstrate that our approach achieves superior performance, highlighting the effectiveness of learning robust, discriminative, and multi-scale features solely from image labels. This work contributes a practical and effective solution for FGSC in remote sensing, reducing the reliance on expensive annotations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AFCN: A Unified Framework for Fine-Grained Ship Classification via Cascaded Attention and Multi-Scale Fusion

  • Nuo Li,
  • Qirong Mao,
  • Meizhu Li

摘要

Fine-grained ship classification (FGSC) in remote sensing imagery is a critical yet challenging task due to significant scale variations, complex backgrounds, high inter-class similarity, and intra-class variability. While advanced deep learning methods have improved performance, a major limitation of many existing approaches is their dependence on costly auxiliary annotations like bounding boxes or part labels, which restricts their scalability and practical deployment. To overcome these issues, we propose a novel FGSC framework for remote sensing imagery that requires only image-level supervision. Our method leverages a powerful convolutional backbone (ConvNeXtV2) enhanced with an integrated cascaded attention mechanism that performs progressive channel-wise refinement and spatial highlighting to improve feature discriminability by dynamically focusing on salient information. Furthermore, it incorporates an innovative weighted bidirectional fusion strategy to effectively combine features from selected high-level network stages, thereby enhancing robustness to scale variations and complex scenes. Extensive experiments on challenging FGSC datasets demonstrate that our approach achieves superior performance, highlighting the effectiveness of learning robust, discriminative, and multi-scale features solely from image labels. This work contributes a practical and effective solution for FGSC in remote sensing, reducing the reliance on expensive annotations.