<p>Infrared small target detection (IRSTD) plays a crucial role in applications such as traffic monitoring systems and maritime rescue. However, existing IRSTD methods face challenges due to their reliance on a single type of data, making them susceptible to noise and deficient in contextual understanding. Additionally, small and limited datasets hinder model generalization and performance in complex scenarios. Previous methods are mostly based on U-Net architectures that are optimized for small-scale data and involve intricate design. These designs often perform well in specific scenarios, but they struggle to generalize effectively in real-world applications. Inspired by leading vision-language models, we propose an MIRSAM (Multimodal Vision-Language Segment Anything Model for Infrared Small Target Detection), the first framework to integrate text modality with image modality for IRSTD in this article. Given the differences in noise and structural information between infrared and natural images, we fine-tune segment anything model (SAM) by designing a contourlet denoising adapter module (CDAM). Integrated into SAM’s image encoder, this module suppresses noise during feature extraction and encoding, enabling efficient adaptation to the infrared domain. To incorporate textual information, we utilize the text encoder of contrastive language-image pre-training (CLIP) to convert text into high-dimensional feature vectors, which then serve as prompts to extract relevant details from the features. In addition, we build the first multimodal IRSTD dataset, IR-TXPair, containing image-text pairs. Experiments on the newly constructed IR-TXPair dataset demonstrate that the proposed MIRSAM outperforms state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MIRSAM: multimodal vision-language segment anything model for infrared small target detection

  • Mingjin Zhang,
  • Qian Xu,
  • Yuchun Wang,
  • Xi Li,
  • Haojuan Yuan

摘要

Infrared small target detection (IRSTD) plays a crucial role in applications such as traffic monitoring systems and maritime rescue. However, existing IRSTD methods face challenges due to their reliance on a single type of data, making them susceptible to noise and deficient in contextual understanding. Additionally, small and limited datasets hinder model generalization and performance in complex scenarios. Previous methods are mostly based on U-Net architectures that are optimized for small-scale data and involve intricate design. These designs often perform well in specific scenarios, but they struggle to generalize effectively in real-world applications. Inspired by leading vision-language models, we propose an MIRSAM (Multimodal Vision-Language Segment Anything Model for Infrared Small Target Detection), the first framework to integrate text modality with image modality for IRSTD in this article. Given the differences in noise and structural information between infrared and natural images, we fine-tune segment anything model (SAM) by designing a contourlet denoising adapter module (CDAM). Integrated into SAM’s image encoder, this module suppresses noise during feature extraction and encoding, enabling efficient adaptation to the infrared domain. To incorporate textual information, we utilize the text encoder of contrastive language-image pre-training (CLIP) to convert text into high-dimensional feature vectors, which then serve as prompts to extract relevant details from the features. In addition, we build the first multimodal IRSTD dataset, IR-TXPair, containing image-text pairs. Experiments on the newly constructed IR-TXPair dataset demonstrate that the proposed MIRSAM outperforms state-of-the-art methods.