Given the increasing emphasis on multimodal data analytics, depth maps have been employed for Salient Object Detection (SOD) task. RGB-D SOD task utilizes the spatial structure information in the depth maps to improve detection accuracy. In this paper, we propose a Transformer-based Depth Optimization Network (DONet) for RGB-D SOD task. A depth feature optimization and integration module (DOIM) is first designed to maximize the auxiliary effect of depth information. In DOIM, high-quality depth information is retained and low-quality information is discarded conversely. Then aiming to comprehensive detail complement, the context supplement modules (CSMs) are configured to absorb features of adjacent layers to refine the features adequately. In addition, for global information exploration, we deploy a location perception guider (LPG) to guide our model to explore the location of salient objects accurately. Based on the wide application of Transformer, the Pyramid Vision Transformer with less computational requirement and equal efficiency is chosen to balance performance and computational cost. Experiments on five widely-used datasets show that the proposed network achieves significant advantage compared to 13 state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Based Depth Optimization Network for RGB-D Salient Object Detection

  • Lu Li,
  • Yanjiao Shi,
  • Jinyu Yang,
  • Qiangqiang Zhou,
  • Qing Zhang,
  • Liu Cui

摘要

Given the increasing emphasis on multimodal data analytics, depth maps have been employed for Salient Object Detection (SOD) task. RGB-D SOD task utilizes the spatial structure information in the depth maps to improve detection accuracy. In this paper, we propose a Transformer-based Depth Optimization Network (DONet) for RGB-D SOD task. A depth feature optimization and integration module (DOIM) is first designed to maximize the auxiliary effect of depth information. In DOIM, high-quality depth information is retained and low-quality information is discarded conversely. Then aiming to comprehensive detail complement, the context supplement modules (CSMs) are configured to absorb features of adjacent layers to refine the features adequately. In addition, for global information exploration, we deploy a location perception guider (LPG) to guide our model to explore the location of salient objects accurately. Based on the wide application of Transformer, the Pyramid Vision Transformer with less computational requirement and equal efficiency is chosen to balance performance and computational cost. Experiments on five widely-used datasets show that the proposed network achieves significant advantage compared to 13 state-of-the-art methods.