Multi-modal semi-supervised semantic segmentation for indoor scenes via adaptive CutMix and contrastive learning
摘要
Semantic segmentation of indoor scenes faces challenges due to uneven lighting, high contrast between light and dark, and cluttered objects, making it difficult to distinguish between foreground and background in RGB images and compromising segmentation effectiveness. To address these issues, we propose ACMatch, an integrated multi-modal semi-supervised semantic segmentation model specifically designed for indoor scenes. ACMatch employs an adaptive CutMix module that leverages labelled data to assist in augmentation of unlabelled data, thereby balancing distributional differences and enhancing segmentation performance under semi-supervision. By combining CNN and Transformer dual branches, ACMatch extracts and fuses features from different modalities, utilising local and global feature alignment to mitigate information loss from downsampling and improve small target segmentation. Additionally, an improved adaptive threshold method is introduced, dynamically adjusting the threshold based on the training stages and set a threshold lower limit to stabilise model training and retain low-confidence yet correct information. A dual-stream contrastive learning module is also constructed to enhance model robustness and generalisation by comparing the feature representations of unlabelled and labelled data. Extensive experiments on the NYUV2 and SUN RGB-D datasets demonstrate that ACMatch initially improves the mIoU index by approximately 4% compared to state-of-the-art semi-supervised methods, significantly enhancing model learning quality and generalisation.