<p>Video action recognition (VAR) aims to analyze dynamic behaviors in videos and achieve semantic understanding. VAR faces challenges such as temporal dynamics, action-scene coupling, and the complexity of human interactions. Existing methods can be categorized into motion-level, event-level, and story-level ones based on spatiotemporal granularity. However, single-modal approaches struggle to capture complex behavioral semantics and human factors. Therefore, in recent years, vision-language models (VLMs) have been introduced into this field, providing new research perspectives for VAR. In this paper, we systematically review spatiotemporal hierarchical methods in VAR and explore how the introduction of large models has advanced the field. Additionally, we propose the concept of “Factor” to identify and integrate key information from both visual and textual modalities, enhancing multimodal alignment. We also summarize various multimodal alignment methods and provide in-depth analysis and insights into future research directions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Video action recognition meets vision-language models exploring human factors in scene interaction: a review

  • Yuping Guo,
  • Hongwei Gao,
  • Jiahui Yu,
  • Jinchao Ge,
  • Meng Han,
  • Zhaojie Ju

摘要

Video action recognition (VAR) aims to analyze dynamic behaviors in videos and achieve semantic understanding. VAR faces challenges such as temporal dynamics, action-scene coupling, and the complexity of human interactions. Existing methods can be categorized into motion-level, event-level, and story-level ones based on spatiotemporal granularity. However, single-modal approaches struggle to capture complex behavioral semantics and human factors. Therefore, in recent years, vision-language models (VLMs) have been introduced into this field, providing new research perspectives for VAR. In this paper, we systematically review spatiotemporal hierarchical methods in VAR and explore how the introduction of large models has advanced the field. Additionally, we propose the concept of “Factor” to identify and integrate key information from both visual and textual modalities, enhancing multimodal alignment. We also summarize various multimodal alignment methods and provide in-depth analysis and insights into future research directions.