Highly accurate dichotomous image segmentation (DIS) [25] has recently attracted research attention, which aims to segment objects from natural images with highly accurate details. However, in previous DIS works, what objects to segment may lack clear definition, and the task remains somewhat elusive. This makes some existing segmentation models under-rated. To address this ambiguity in DIS, this paper introduces a modified task, called prior mask-guided DIS (PMG-DIS), where polygon masks covering desired objects are annotated by users. We then propose a Prior Fusion Module (PFM) specifically designed to exploit user-input auxiliary positional information and further develop a specialized model named Bi-Stream Refinement Network (BSRNet), which has two stages: identification and refinement. For identification, we utilize a “Scale Reintegration Decoder” (SRD) derived from dual streams to extract and fuse different types/levels of features to initially identify and segment target objects. During refinement, we design a Cross-level Aggregation Module (CAM) used to refine coarse feature and enhance object details with refinement blocks. Experimental results are conducted on DIS5K dataset. Our BSRNet outperforms 10 state-of-the-art models on the PMG-DIS task by notable margins. Besides, we show that all previous models can be improved when being fed with prior masks. Our code and dataset will be publicly available.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prior Mask-Guided Highly Accurate Dichotomous Image Segmentation

  • Shanfeng Zhou,
  • Bo Yuan,
  • Keren Fu,
  • Hailun Zhang,
  • Qijun Zhao

摘要

Highly accurate dichotomous image segmentation (DIS) [25] has recently attracted research attention, which aims to segment objects from natural images with highly accurate details. However, in previous DIS works, what objects to segment may lack clear definition, and the task remains somewhat elusive. This makes some existing segmentation models under-rated. To address this ambiguity in DIS, this paper introduces a modified task, called prior mask-guided DIS (PMG-DIS), where polygon masks covering desired objects are annotated by users. We then propose a Prior Fusion Module (PFM) specifically designed to exploit user-input auxiliary positional information and further develop a specialized model named Bi-Stream Refinement Network (BSRNet), which has two stages: identification and refinement. For identification, we utilize a “Scale Reintegration Decoder” (SRD) derived from dual streams to extract and fuse different types/levels of features to initially identify and segment target objects. During refinement, we design a Cross-level Aggregation Module (CAM) used to refine coarse feature and enhance object details with refinement blocks. Experimental results are conducted on DIS5K dataset. Our BSRNet outperforms 10 state-of-the-art models on the PMG-DIS task by notable margins. Besides, we show that all previous models can be improved when being fed with prior masks. Our code and dataset will be publicly available.