错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SANet: similarity aggregation and semantic fusion for few-shot semantic segmentation

  • Minrui Ye,
  • Tao Zhang

摘要

Few-shot semantic segmentation (FSS) methods based on meta-learning strategies have shown promise in extracting instance knowledge from support set to infer pixel-wise labels in query set. However, a key challenge in FSS is addressing spatial inconsistency between query image and support image due to intra-class difference and inter-class similarity. Moreover, existing FSS methods often rely on multiple decoding methods for differentiated pixel-wise matching, leading to semantic inconsistency. To tackle these issues, we propose a similarity aggregation network (SANet), which effectively explores visual correspondence between support and query features while aligning semantic dimensions. Specifically, SANet introduces a mask attention module (MAM) to capture spatial relations between non-local attention features from support features and query features. Additionally, a similarity aggregation module (SAM) is proposed, which utilizes the multi-head attention mechanism and combines prior mask to calculate the aggregation similarity between each query pixel and all supporting pixels, thereby focusing the network on foreground areas. Finally, a feature fusion module (FFM) is used to adaptively fuse features at multiple scales and channels for accurate prediction. Extensive experiments on PASCAL-5i and COCO-20i demonstrate the efficiency and competitiveness of SANet.