Recently, self-supervised methods based on self-supervised transformer features have demonstrated promising results in unsupervised object localization. However, obtaining exceptional semantic results remains a formidable challenge. The current approaches heavily rely on the similarity between patch-level features within an image, while lacking supervision from image-level information. Meanwhile, the Segment Anything Model (SAM) has demonstrated remarkable class-agnostic segmentation capabilities for arbitrary objects in images with sparse prompts like points. In this work, we propose SelfLoc, a simple yet effective self-supervised object localization method via integration with self-prompt SAM. Specifically, a self-prompt generator is designed to automatically generate sparse prompts based on an image’s self-attention map. Simultaneously, an image-wise integration module is developed to enhance the coarse mask obtained from self-supervised features by leveraging the fine-grained segmentation results of SAM. Extensive experimental results demonstrate that the proposed method not only achieves state-of-the-art performance in unsupervised saliency detection and object discovery tasks, but also sets a new benchmark in unsupervised camouflaged object segmentation. The source code will be publicly available at https://github.com/Rogersiy/SelfLoc .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SelfLoc: High Quality Unsupervised Object Localization with Self-Prompt SAM

  • Jiaheng Zhang,
  • Xiandong Wang,
  • Conghui Li,
  • Longyi Chen,
  • Shengke Wang

摘要

Recently, self-supervised methods based on self-supervised transformer features have demonstrated promising results in unsupervised object localization. However, obtaining exceptional semantic results remains a formidable challenge. The current approaches heavily rely on the similarity between patch-level features within an image, while lacking supervision from image-level information. Meanwhile, the Segment Anything Model (SAM) has demonstrated remarkable class-agnostic segmentation capabilities for arbitrary objects in images with sparse prompts like points. In this work, we propose SelfLoc, a simple yet effective self-supervised object localization method via integration with self-prompt SAM. Specifically, a self-prompt generator is designed to automatically generate sparse prompts based on an image’s self-attention map. Simultaneously, an image-wise integration module is developed to enhance the coarse mask obtained from self-supervised features by leveraging the fine-grained segmentation results of SAM. Extensive experimental results demonstrate that the proposed method not only achieves state-of-the-art performance in unsupervised saliency detection and object discovery tasks, but also sets a new benchmark in unsupervised camouflaged object segmentation. The source code will be publicly available at https://github.com/Rogersiy/SelfLoc .