WS-GCA: A Synergistic Framework for Precise Semantic Segmentation with Comprehensive Supervision
摘要
Semantic segmentation is a fundamental task in computer vision that entails classifying each pixel of an image into predefined categories. Despite significant advancements in deep learning, obtaining accurately labeled datasets remains a costly and labor-intensive process. This research aims to mitigate the need for extensive, precise tags by exploring Weakly Supervised Semantic Segmentation (WSSS), which seeks to achieve accurate pixel-level classification with minimal supervision. We introduce WS-GCA, a novel unified framework that synergistically combines the Gaussian Mixture Model (GMM), Label Cohesion Loss (LC Loss), and self-attention mechanism to enhance segmentation quality. The WS-GCA framework models the distribution of weak labels using a mixed Gaussian distribution, amalgamates global and local feature information to substantially boost model prediction accuracy, incorporates LC Loss to improve spatial consistency in segmentation, and employs a self-attention mechanism to enhance feature extraction efficiency. Experimental results on the Pascal and Cityscapes datasets demonstrate the WS-GCA framework’s ability to generate superior segmentation results from initially weak labels. The proposed framework increases the mean Intersection over Union (mIoU) by 2.2% compared to baseline models, significantly reducing category mispredictions and advancing the state of the art in the segmentation of large-area objects with minimal supervision.