<p>Detecting group activities in surveillance videos poses significant challenges due to the complex interactions among multiple actors. Existing methods often employ RoIAlign for feature extraction, yielding coarse actor representations that are prone to background noise interference. Furthermore, these approaches typically neglect language knowledge, which restricts their ability to achieve robust semantic understanding. To overcome these shortcomings, we propose the Visual-Language Collaborative Multimodal Transformer Network (VLCMTN), a novel framework that enhances group activity detection by effectively combining visual and language information. Our method introduces two primary innovations: (1) We adopt an off-the-shelf open-world panoptic segmentation model with box prompts to produce precise actor embeddings, effectively minimizing background interference. (2) We provide a new perspective on group activity detection by emphasizing the semantic information of language rather than merely mapping visual features to labels. Specifically, off-the-shelf Video Multimodal Large Language Model (MLLM) is adopted to generate detailed video descriptions. Moreover, we propose a MLLM knowledge-enhanced multimodal transformer that jointly processes visual and language cues. To ensure relevance, we incorporate a Top-k sparse aggregation mechanism to select the most pertinent language information. By modeling fine-grained intra- and inter-modal relationships using language-guided knowledge, extensive experiments on public benchmarks demonstrate that our approach achieves state-of-the-art results.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual-language collaborative multimodal transformer network for group activity detection in surveillance videos

  • Fudong Nian,
  • Weijie Lu,
  • Jun Wang,
  • Chengqian Li,
  • Yun Fu,
  • Zhize Wu

摘要

Detecting group activities in surveillance videos poses significant challenges due to the complex interactions among multiple actors. Existing methods often employ RoIAlign for feature extraction, yielding coarse actor representations that are prone to background noise interference. Furthermore, these approaches typically neglect language knowledge, which restricts their ability to achieve robust semantic understanding. To overcome these shortcomings, we propose the Visual-Language Collaborative Multimodal Transformer Network (VLCMTN), a novel framework that enhances group activity detection by effectively combining visual and language information. Our method introduces two primary innovations: (1) We adopt an off-the-shelf open-world panoptic segmentation model with box prompts to produce precise actor embeddings, effectively minimizing background interference. (2) We provide a new perspective on group activity detection by emphasizing the semantic information of language rather than merely mapping visual features to labels. Specifically, off-the-shelf Video Multimodal Large Language Model (MLLM) is adopted to generate detailed video descriptions. Moreover, we propose a MLLM knowledge-enhanced multimodal transformer that jointly processes visual and language cues. To ensure relevance, we incorporate a Top-k sparse aggregation mechanism to select the most pertinent language information. By modeling fine-grained intra- and inter-modal relationships using language-guided knowledge, extensive experiments on public benchmarks demonstrate that our approach achieves state-of-the-art results.