STNet: a lightweight spectral transform framework for salient object detection
摘要
Salient object detection (SOD) is a key research direction in the field of computer vision and has attracted extensive attention from scholars in this field. Although deep learning has made significant progress, two key bottlenecks remain: (1) Existing methods fail to reconcile precise edge detection in low-level features with semantic coherence in high-level features, resulting in compromised boundary integrity for complex-shaped objects. (2) Conventional architectures exhibit inherent sensitivity to appearance variations due to their spatial domain limitations, and lack frequency-adaptive robustness. To address these issues, this paper proposes a novel lightweight spectral transform framework (STNet) for SOD. First, a multi-feature fusion network is introduced as a baseline model for saliency inference. An edge-guiding module is used to extract precise boundaries via differential pooling, and a semantic fusion module aligns cross-level features via dynamic dilated convolutions , both of which are integrated into this network. The objective of this integration is to efficiently aggregate fine-grained visual features and abstract semantic information while guaranteeing feature space consistency. Second, to enhance the robustness of salient information, we incorporate a spectral transform module that combines spatial and frequency domain features. This module highlights target details in multi-frequency domains and improves saliency prediction. Finally, a simple yet effective optimized loss function is designed to refine the saliency predictions. Extensive experiments confirm that the proposed STNet outperforms competing methods and is capable of accurately detecting large targets, multiple targets, and objects in both simple and complex scenarios. It achieves MAEs (