<p>Video scene graph generation (VidSGG) aims to predict visual relation triplets in videos, which is a key step toward a deeper comprehension of video scenes. To enhance applicability in real-world scenarios, open-vocabulary settings have been explored in recent VidSGG frameworks. However, the heavy reliance on aligned visual and textual representations from pre-trained vision-language models (VLMs) limits the in-depth understanding of visual relationships, as they fail to comprehend compositional scene relationships. To address this, we propose a novel open-vocabulary VidSGG framework named semantic-unified cross-modal learning (SUCML), which leverages the exceptional capabilities of large language models (LLMs) in visual understanding and semantic reasoning to achieve robust visual relation prediction. Specifically, we incorporate rich knowledge and open-vocabulary capabilities into our framework and design a cross-modal adapter to facilitate the semantic-unified cross-modal representation learning process. We then utilize the impressive semantic understanding and reasoning capabilities of LLMs, providing predefined textual instructions along with the learned semantic-unified visual and text tokens as inputs to the pre-trained LLM for robust relation prediction. Extensive experimental results on two public datasets demonstrate that SUCML significantly outperforms existing methods, showcasing the promising potential of LLMs for high-level semantic reasoning and scene understanding.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning semantic-unified cross-modal representations for open-vocabulary video scene graph generation

  • Yufan Hu,
  • Fang Zhang,
  • Ran Wei,
  • Junling Gao

摘要

Video scene graph generation (VidSGG) aims to predict visual relation triplets in videos, which is a key step toward a deeper comprehension of video scenes. To enhance applicability in real-world scenarios, open-vocabulary settings have been explored in recent VidSGG frameworks. However, the heavy reliance on aligned visual and textual representations from pre-trained vision-language models (VLMs) limits the in-depth understanding of visual relationships, as they fail to comprehend compositional scene relationships. To address this, we propose a novel open-vocabulary VidSGG framework named semantic-unified cross-modal learning (SUCML), which leverages the exceptional capabilities of large language models (LLMs) in visual understanding and semantic reasoning to achieve robust visual relation prediction. Specifically, we incorporate rich knowledge and open-vocabulary capabilities into our framework and design a cross-modal adapter to facilitate the semantic-unified cross-modal representation learning process. We then utilize the impressive semantic understanding and reasoning capabilities of LLMs, providing predefined textual instructions along with the learned semantic-unified visual and text tokens as inputs to the pre-trained LLM for robust relation prediction. Extensive experimental results on two public datasets demonstrate that SUCML significantly outperforms existing methods, showcasing the promising potential of LLMs for high-level semantic reasoning and scene understanding.