Dynamic Self-supervision for Scribble-Based Semantic Segmentation via Stripe Pooling and Coordinate Attention
摘要
Semantic segmentation is one of the three fundamental tasks in computer vision. Recent advances have achieved significant success with precise, pixel-wise annotated training. However, collecting such dense annotations always requires laborious manual work. Sparse annotations, on the other hand, are cost-effective and contain the minimal necessary class and location information. Therefore, semantic segmentation with sparse annotations offers high research potential in balancing information and cost. Nevertheless, scribble labels, as a form of sparse annotation, present challenges such as insufficient annotation information and inadequate regional information. To address these issues, this paper proposes a dynamic self-supervised semantic segmentation model. Building upon DeeplabV3+, the weakly-supervised semantic segmentation model achieves online dynamic self-supervision. This model is designed to perform adaptive supervision using scribble labels to complement the training of the primary segmentation model. For enhanced utilization of the auxiliary branch’s clustering impact, the predictive main branch of the model integrates a stripe pooling module and a coordinate attention mechanism module, both of which are subject to joint supervision. Experiments on the PASCAL VOC 2012 dataset demonstrate that this model achieves high accuracy with scribble labels, obtaining a mean Intersection over Union (mIoU) of 79.0%.