<p>Saliency prediction (SP) estimates human gaze fixation in scenes, guided by visual attention mechanisms. While deep learning approaches have made substantial progress in SP, many overlook the inherent scene information presented in images. Further exploration is needed to generalize SP models across broader scene types and investigate how SP models behave differently across scenes, thereby advancing our understanding of scene knowledge’s impact on SP performance and human visual attention. This study introduces a scene-specific SP framework that incorporates scene labels predicted by a transferred CLIP-based classifier with multi-layer perceptron enhancement, trained on the CAT2000 dataset comprising 20 scene types. The scene classifier achieves high accuracy (averaging 88%) for 20-scene classification, providing reliable labels for subsequent scene-specific SP. In the SP phase, we employ transfer learning with DINet, training separate SP models for each scene type and a general SP model trained on images from all 20 scene types for comparison. Results show that the scene-specific SP framework consistently outperforms the general model across most scenes, with average improvements of 0.36% in AUC, 6.07% in NSS, and 4.66% in CC. Higher NSS and CC gains indicate that the scene-specific models effectively capture more accurate saliency position and distribution information from each specific scene that aligns with human gaze patterns. Moreover, the proposed framework is model-agnostic, ensuring compatibility with various SP models. Our findings highlight the cognitive importance of incorporating prior scene knowledge for precise SP and deepen the understanding of visual attention mechanisms across diverse environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Scene-Specific Saliency Prediction Incorporating Scene Information

  • Shuning Han,
  • Zhe Sun,
  • Takashi Michikawa,
  • Cesar F. Caiafa,
  • Jordi Solé-Casals,
  • Hideo Yokota

摘要

Saliency prediction (SP) estimates human gaze fixation in scenes, guided by visual attention mechanisms. While deep learning approaches have made substantial progress in SP, many overlook the inherent scene information presented in images. Further exploration is needed to generalize SP models across broader scene types and investigate how SP models behave differently across scenes, thereby advancing our understanding of scene knowledge’s impact on SP performance and human visual attention. This study introduces a scene-specific SP framework that incorporates scene labels predicted by a transferred CLIP-based classifier with multi-layer perceptron enhancement, trained on the CAT2000 dataset comprising 20 scene types. The scene classifier achieves high accuracy (averaging 88%) for 20-scene classification, providing reliable labels for subsequent scene-specific SP. In the SP phase, we employ transfer learning with DINet, training separate SP models for each scene type and a general SP model trained on images from all 20 scene types for comparison. Results show that the scene-specific SP framework consistently outperforms the general model across most scenes, with average improvements of 0.36% in AUC, 6.07% in NSS, and 4.66% in CC. Higher NSS and CC gains indicate that the scene-specific models effectively capture more accurate saliency position and distribution information from each specific scene that aligns with human gaze patterns. Moreover, the proposed framework is model-agnostic, ensuring compatibility with various SP models. Our findings highlight the cognitive importance of incorporating prior scene knowledge for precise SP and deepen the understanding of visual attention mechanisms across diverse environments.