Selective Targeting for Enhanced Salient Object Detection
摘要
In salient object detection, many current methods apply the same processing to features across different scales and stages, neglecting the variations in these features and their impact on prediction. High-level features offer rich semantic information but lack precise spatial details, while low-level features provide more detail but often include background noise, complicating accurate saliency target localization and affecting overall model performance. This study introduces a new network architecture, SANet, incorporating a dynamic token selector (DTS) to selectively extract key features and reduce the influence of interfering features. To enhance saliency map generation, we propose a target-focused attention mechanism (TFA), combining query-centered sliding window attention with pooling attention to improve the model’s handling of multi-scale features. This design enhances feature expressiveness and the accuracy of salient object detection. Additionally, we introduce a joint supervision strategy to refine features at each level. Experiments on five challenging datasets show SANet’s superiority over current state-of-the-art methods in salient object detection.