In the field of computer vision, semantic segmentation has consistently garnered significant attention. To enhance the performance of diffusion models in the domain of weak supervision, this paper proposes an improved weakly supervised semantic segmentation method, termed WSSS-DM+. The WSSS-DM+ method incorporates a DC module into the original model architecture based on the cross-attention mechanism and employs a variety of model optimization strategies, which collectively enhance the quality of the generated masks. Additionally, this method more accurately establishes associations between text and image regions, thereby achieving a visual explanation of the text-image diffusion model. Experimental investigations have demonstrated the generalizability of this method across different texts. Furthermore, experimental results indicate that, compared to the WSSS-DM method, the new approach effectively addresses the issue of coarse masks, with the mIoU and mAcc metrics improving by 4.7 and 5.8, respectively, thus confirming the effectiveness of the proposed enhancements.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

WSSS-DM+: Advancing Improvements in Weakly Supervised Semantic Segmentation Using Diffusion Models

  • Zhaofeng Niu,
  • Xingfu Cheng,
  • Chuanguo Shen,
  • Bowen Wang,
  • Guangshun Li,
  • Liangzhi Li

摘要

In the field of computer vision, semantic segmentation has consistently garnered significant attention. To enhance the performance of diffusion models in the domain of weak supervision, this paper proposes an improved weakly supervised semantic segmentation method, termed WSSS-DM+. The WSSS-DM+ method incorporates a DC module into the original model architecture based on the cross-attention mechanism and employs a variety of model optimization strategies, which collectively enhance the quality of the generated masks. Additionally, this method more accurately establishes associations between text and image regions, thereby achieving a visual explanation of the text-image diffusion model. Experimental investigations have demonstrated the generalizability of this method across different texts. Furthermore, experimental results indicate that, compared to the WSSS-DM method, the new approach effectively addresses the issue of coarse masks, with the mIoU and mAcc metrics improving by 4.7 and 5.8, respectively, thus confirming the effectiveness of the proposed enhancements.