Although vision-language knowledge exploitation achieves great success in some visual tasks, its effectiveness has not been well investigated in the field of video saliency prediction (VSP) so far. Moreover, it would be suboptimal to simply encode language information into a vision model. Our motivation is to explore the interaction between vision and language knowledge in the VSP task. Thus, in the work, we propose a novel VSP model based on a Transformer-based Vision-Language Interaction Network (TVLI-Net), pioneering the integration of visual and linguistic data in VSP. Specifically, our model first leverages pre-trained encoders to consolidate respectively semantic information from both the visual and linguistic realms. A Word Feature Weighting (WFW) module is then introduced to adaptively adjust the relative importance of each word to visual saliency. After a slight vision-language interaction is established via WFW module, Vision-Language Interaction modules based on Cross Concatenation (VLI-CC) are introduced to further enhance the interaction. In our decoder architecture, we present Language-Guided Gate Attention mechanism (LGGA) modules, channeling visual attention towards salient regions guided by linguistic global information. Extensive experiments on benchmark datasets such as DHF1K, Hollywood-2, and UCF Sports clearly validate the superiority of our model over existing state-of-the-art models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision-Language Knowledge Exploration for Video Saliency Prediction

  • Fei Zhou,
  • Baitao Huang,
  • Guoping Qiu

摘要

Although vision-language knowledge exploitation achieves great success in some visual tasks, its effectiveness has not been well investigated in the field of video saliency prediction (VSP) so far. Moreover, it would be suboptimal to simply encode language information into a vision model. Our motivation is to explore the interaction between vision and language knowledge in the VSP task. Thus, in the work, we propose a novel VSP model based on a Transformer-based Vision-Language Interaction Network (TVLI-Net), pioneering the integration of visual and linguistic data in VSP. Specifically, our model first leverages pre-trained encoders to consolidate respectively semantic information from both the visual and linguistic realms. A Word Feature Weighting (WFW) module is then introduced to adaptively adjust the relative importance of each word to visual saliency. After a slight vision-language interaction is established via WFW module, Vision-Language Interaction modules based on Cross Concatenation (VLI-CC) are introduced to further enhance the interaction. In our decoder architecture, we present Language-Guided Gate Attention mechanism (LGGA) modules, channeling visual attention towards salient regions guided by linguistic global information. Extensive experiments on benchmark datasets such as DHF1K, Hollywood-2, and UCF Sports clearly validate the superiority of our model over existing state-of-the-art models.