Asymmetric Dual-Stream Networks for Lightweight RGB-D Salient Object Detection
摘要
Integrating image and depth information for RGB-D salient object detection has become a research hotspot in the field of saliency detection. Balancing the efficiency and performance of salient object detection models under resource constraints is a key challenge. To address this, this paper proposes an asymmetric lightweight network suitable for real-time RGB-D salient object detection tasks. The network reduces the number of network parameters by designing different lightweight feature extraction networks for different input modalities. Additionally, a multi-modal feature enhancement fusion module is designed to effectively fuse multi-modal features while compensating for the information loss caused by the lightweight backbone network. Moreover, this paper utilizes a global context module for dense decoding, aggregating local and global information of multi-scale features without significantly increasing computational complexity. The experimental results on five benchmarks show that the proposed lightweight RGB-D salient object detection network not only outperforms most mainstream models quantitatively and qualitatively, but also significantly outperforms other models in terms of efficiency, only with a parameter count of 5.1 million and a computational load of 0.77 gigaflops. This achievement validates the proposed method’s ability to achieve lightweight salient object detection while maintaining high efficiency.