<p>Weakly supervised semantic segmentation aims to achieve dense segmentation using minimal annotations. However, standard class activation maps (CAMs) typically exhibit sparse and incomplete activations. While existing solutions often expand these activations through parameter-coupled optimization, such adaptation may introduce dataset-specific biases and additional optimization costs, especially under limited image-level supervision. To address this, we introduce self-visual prompting (SVP), a training-free CAM generation paradigm that employs a recursive self-bootstrapping strategy to shift optimization from model parameters to the input context. By utilizing background blurring as a negative visual prompt, SVP empowers frozen vision-language models to autonomously correct their attentional focus. Specifically, our Synergistic CAM extraction module acts as a prompt interpreter, decoding reliable semantic anchors by synergizing dynamic gradient responses with static structural consistency. Guided by this semantic prior, the hierarchical prompt generator constructs spatial negative visual prompts by selectively blurring distractor regions while strictly maintaining object boundaries. Through the recursive interplay of these two modules, our approach progressively refines the quality of CAMs. Without any parameter updates during the CAM generation stage, SVP establishes a new state-of-the-art, achieving final segmentation mIoUs of 75.2% on PASCAL VOC dataset and 47.5% on MS COCO dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-visual prompting for training-free recursive weakly supervised semantic segmentation

  • Congwei Zhang,
  • Zhibin Quan,
  • Yuncong Yao,
  • Wankou Yang

摘要

Weakly supervised semantic segmentation aims to achieve dense segmentation using minimal annotations. However, standard class activation maps (CAMs) typically exhibit sparse and incomplete activations. While existing solutions often expand these activations through parameter-coupled optimization, such adaptation may introduce dataset-specific biases and additional optimization costs, especially under limited image-level supervision. To address this, we introduce self-visual prompting (SVP), a training-free CAM generation paradigm that employs a recursive self-bootstrapping strategy to shift optimization from model parameters to the input context. By utilizing background blurring as a negative visual prompt, SVP empowers frozen vision-language models to autonomously correct their attentional focus. Specifically, our Synergistic CAM extraction module acts as a prompt interpreter, decoding reliable semantic anchors by synergizing dynamic gradient responses with static structural consistency. Guided by this semantic prior, the hierarchical prompt generator constructs spatial negative visual prompts by selectively blurring distractor regions while strictly maintaining object boundaries. Through the recursive interplay of these two modules, our approach progressively refines the quality of CAMs. Without any parameter updates during the CAM generation stage, SVP establishes a new state-of-the-art, achieving final segmentation mIoUs of 75.2% on PASCAL VOC dataset and 47.5% on MS COCO dataset.