<p>RGB-D salient object segmentation exploits the complementary characteristics of RGB appearance cues and depth geometric information, yet its performance is often hindered by noisy and spatially unreliable depth data as well as insufficient global-to-local refinement mechanisms. To address these issues, this paper proposes a depth reliability-aware dual-stage RGB-D salient object segmentation framework. The proposed method introduces a multi-level cross-modal feature fusion strategy that adaptively balances RGB semantic information and depth structural cues via modality-aware reweighting and residual semantic anchoring. On this basis, a coarse-to-fine segmentation paradigm is employed, where a lightweight decoder first generates a global coarse saliency mask that subsequently guides a mask-based refinement process to progressively enhance object boundaries and structural details under top-down semantic constraints. Extensive experiments on public RGB-D saliency benchmarks demonstrate that the proposed approach achieves strong structural consistency and competitive segmentation performance compared with recent state-of-the-art methods. In particular, the proposed framework shows advantages in preserving global object structure and maintaining stable pixel-level prediction under complex scenes with unreliable depth information.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Depth reliability-aware dual-stage network for RGB-D salient object segmentation

  • Hongjie He,
  • Haolin Huang,
  • Yizheng Wang

摘要

RGB-D salient object segmentation exploits the complementary characteristics of RGB appearance cues and depth geometric information, yet its performance is often hindered by noisy and spatially unreliable depth data as well as insufficient global-to-local refinement mechanisms. To address these issues, this paper proposes a depth reliability-aware dual-stage RGB-D salient object segmentation framework. The proposed method introduces a multi-level cross-modal feature fusion strategy that adaptively balances RGB semantic information and depth structural cues via modality-aware reweighting and residual semantic anchoring. On this basis, a coarse-to-fine segmentation paradigm is employed, where a lightweight decoder first generates a global coarse saliency mask that subsequently guides a mask-based refinement process to progressively enhance object boundaries and structural details under top-down semantic constraints. Extensive experiments on public RGB-D saliency benchmarks demonstrate that the proposed approach achieves strong structural consistency and competitive segmentation performance compared with recent state-of-the-art methods. In particular, the proposed framework shows advantages in preserving global object structure and maintaining stable pixel-level prediction under complex scenes with unreliable depth information.