MambaSOD-Lite: Efficient RGB-D Saliency Detection via State Space Modeling
摘要
RGB-D Salient Object Detection (SOD) aims to accurately identify perceptually dominant regions in visual content by leveraging complementary cues from color and depth modalities. Existing approaches often rely on attention-intensive fusion designs or transformer-based architectures, resulting in increased computational overhead and limited adaptability to varying depth quality. We propose MambaSOD-Lite, an expressive and lightweight framework that combines dual state-space encoders with a newly introduced Cross-Modal Mamba (CMM) module for efficient feature fusion. The architecture integrates depthwise separable convolutions, channel attention mechanisms, and linear-complexity scanning to retain global context while optimizing hardware efficiency, while a multi-scale refinement decoder enhances boundary precision through hierarchical feature aggregation. Extensive experiments on diverse RGB-D benchmarks validate the effectiveness of our method, achieving strong detection performance and real-time throughput—setting a new standard for accurate and deployable RGB-D SOD.