<p>The self-attention mechanism of vision transformer has demonstrated remarkable superiority in fully-supervised instance segmentation. However, when supervision is constrained to a single point per instance, self-attention struggles to capture semantic variations across object parts, particularly for objects with substantial deformations and diverse appearances. In this study, we propose discriminatively matched part tokens (DMPT), to endow self-attention the capability of handling significant semantic variation for pointly supervised instance segmentation. DMPT first allocates a token for each object part by finding a semantic extreme point, and then introduces part classifiers with deformable constraint to re-estimate part tokens which are utilized to guide and enhance the fine-grained localization capability of the self-attention mechanism. Through iterative optimization, DMPT matches the most discriminative tokens which facilitate capturing fine-grained part semantics and activating full object extent. Extensive experiments on PASCAL VOC and MS COCO segmentation datasets show that DMPT respectively outperforms the state-of-the-art pointly supervised method by 2.0% mAP<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2547_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="13" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{50}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>50</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> and 1.6% AP. When integrated with Segment Anything Model (SAM), DPMT achieves further performance gains, demonstrating the efficacy of simple point-based prompts in enhancing the effectiveness of point-supervised segmentation. The code is available at <a href="https://github.com/guozonghao96/DMPT">https://github.com/guozonghao96/DMPT</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation

  • Zonghao Guo,
  • Fang Wan,
  • Mingxiang Liao,
  • Yidan Zhang,
  • Qixiang Ye

摘要

The self-attention mechanism of vision transformer has demonstrated remarkable superiority in fully-supervised instance segmentation. However, when supervision is constrained to a single point per instance, self-attention struggles to capture semantic variations across object parts, particularly for objects with substantial deformations and diverse appearances. In this study, we propose discriminatively matched part tokens (DMPT), to endow self-attention the capability of handling significant semantic variation for pointly supervised instance segmentation. DMPT first allocates a token for each object part by finding a semantic extreme point, and then introduces part classifiers with deformable constraint to re-estimate part tokens which are utilized to guide and enhance the fine-grained localization capability of the self-attention mechanism. Through iterative optimization, DMPT matches the most discriminative tokens which facilitate capturing fine-grained part semantics and activating full object extent. Extensive experiments on PASCAL VOC and MS COCO segmentation datasets show that DMPT respectively outperforms the state-of-the-art pointly supervised method by 2.0% mAP \(_{50}\) 50 and 1.6% AP. When integrated with Segment Anything Model (SAM), DPMT achieves further performance gains, demonstrating the efficacy of simple point-based prompts in enhancing the effectiveness of point-supervised segmentation. The code is available at https://github.com/guozonghao96/DMPT