TIINet: A Three-Stage Interactive Integration Network for RGB-D Salient Object Detection
摘要
Current research on salient object detection predominantly focuses on optimizing intermediate features to enhance the utilization of multimodal information. However, most models overlook the differences between features at various levels during the encoding and decoding stages, leading to insufficient feature utilization and consequently limiting model performance. To address this issue, we propose the Three-stage Interactive Integration Network (TIINet). In the encoding phase, we introduce a three-stage feature optimization module to process RGB and depth features at different stages, enhancing the overall representation. In the decoding stage, we propose a three-stage feature aggregation module, which effectively fuses multimodal features by integrating features at different levels, thereby improving model performance. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) methods, achieving significant performance improvements across five benchmark datasets: STERE, NJU2K, NLPR, SIP, and SSD.