Video action recognition meets vision-language models exploring human factors in scene interaction: a review
摘要
Video action recognition (VAR) aims to analyze dynamic behaviors in videos and achieve semantic understanding. VAR faces challenges such as temporal dynamics, action-scene coupling, and the complexity of human interactions. Existing methods can be categorized into motion-level, event-level, and story-level ones based on spatiotemporal granularity. However, single-modal approaches struggle to capture complex behavioral semantics and human factors. Therefore, in recent years, vision-language models (VLMs) have been introduced into this field, providing new research perspectives for VAR. In this paper, we systematically review spatiotemporal hierarchical methods in VAR and explore how the introduction of large models has advanced the field. Additionally, we propose the concept of “Factor” to identify and integrate key information from both visual and textual modalities, enhancing multimodal alignment. We also summarize various multimodal alignment methods and provide in-depth analysis and insights into future research directions.