<p>Detecting small objects in high-resolution images remains a challenging task due to limited discriminative cues, occlusions, and large intra-class variations. In this paper, we propose an enhanced Transformer-based detection framework to address these challenges. Our approach employs a ResNet backbone for multi-scale feature extraction and introduces a novel adaptive contrast enhancement module, which highlights low-contrast regions by combining the original high-resolution input image with its corresponding high-resolution feature map. Additionally, we propose a contrast-attentive fusion mechanism that leverages deformable downsampling to integrate contrast-enhanced information with high-level features, thereby emphasizing small-object regions. We further incorporate attention-based intra-scale feature interaction and CNN-based cross-scale feature fusion modules to refine multi-scale representations. Finally, we design a small-object-aware uncertainty-minimal query selection mechanism that preserves essential queries for small objects while reducing computational overhead. Experimental results on VisDrone and MS-COCO demonstrate that our model significantly improves small object detection performance in high-resolution scenarios compared to recent state-of-the-art methods, confirming the effectiveness of our proposed modules.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A contrast-aware and uncertainty-optimized transformer model for small object detection in high-resolution images

  • Nguyen Hoanh,
  • Tran Vu Pham

摘要

Detecting small objects in high-resolution images remains a challenging task due to limited discriminative cues, occlusions, and large intra-class variations. In this paper, we propose an enhanced Transformer-based detection framework to address these challenges. Our approach employs a ResNet backbone for multi-scale feature extraction and introduces a novel adaptive contrast enhancement module, which highlights low-contrast regions by combining the original high-resolution input image with its corresponding high-resolution feature map. Additionally, we propose a contrast-attentive fusion mechanism that leverages deformable downsampling to integrate contrast-enhanced information with high-level features, thereby emphasizing small-object regions. We further incorporate attention-based intra-scale feature interaction and CNN-based cross-scale feature fusion modules to refine multi-scale representations. Finally, we design a small-object-aware uncertainty-minimal query selection mechanism that preserves essential queries for small objects while reducing computational overhead. Experimental results on VisDrone and MS-COCO demonstrate that our model significantly improves small object detection performance in high-resolution scenarios compared to recent state-of-the-art methods, confirming the effectiveness of our proposed modules.