Bi-directional Interaction and Dense Aggregation Network for RGB-D Salient Object Detection
摘要
RGB-D salient object detection (SOD) which aims to detect the prominent regions in figures has attracted much attention recently. It jointly models the RGB and depth information. However, existing methods explore cross-modality information from RGB images and depth maps without considering the potential coupling correlation between them. This may lead to insufficient information learning of these two modalities and even bring conflict due to their de-coupled representations. Thus, in this paper, we propose a novel framework called Bi-directional Interaction and Dense Aggregation Network (BIDANet) for RGB-D salient object detection. Firstly, we carefully design the depth-guided enhancement (DGE) and RGB-induced style transfer (RST) to allow the depth map and RGB image to learn information from each other through the bi-directional interaction network. Secondly, we adopt an adaptive cross-modal fusion (ACF) to flexibly integrate these learned multi-modal features. Last, we propose a dense aggregation network (DAN) to effectively aggregate cross-stage outcomes and generate accurate saliency prediction. Extensive experiments on 5 widely-used datasets demonstrate that our proposed BIDANet achieves superior performance compared with 14 state-of-the-art methods.