This paper introduces the Character Identifying Video Language Alignment Network (CIVLAN), a labor-efficient approach for weakly-supervised Video-Subtitle Moment Retrieval (VSMR). CIVLAN efficiently localizes pertinent temporal moment in untrimmed videos or subtitles by responding to natural language queries. It addresses three key limitations of previous methods: the extensive need for fully supervised training with labor-intensive annotations, underutilization of auxiliary modalities like subtitles and audio, and the neglection of character identification in dialogues. CIVLAN comprises two main components: the Modality Alignment Network (MAN), which aligns queries to video/subtitle content through contrastive loss minimization, and the Character Identifying Network (CIN), which associates on-screen characters with their references in subtitles and queries. This method eliminates the need for ground-truth temporal moment, significantly reducing manual annotation efforts. Our experiments demonstrate that CIVLAN surpasses previous state-of-the-art methods on the TVR benchmark dataset under weakly-supervised conditions. The paper also includes comprehensive ablation studies and qualitative analyses to showcase the effectiveness of CIVLAN’s components.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Character Identifying Video Language Alignment Network for Weakly-Supervised Video-Subtitle Moment Retrieval

  • Donghoon Lee,
  • Jun Yeop Shim,
  • Sun-Jae Yoon,
  • Chang D. Yoo

摘要

This paper introduces the Character Identifying Video Language Alignment Network (CIVLAN), a labor-efficient approach for weakly-supervised Video-Subtitle Moment Retrieval (VSMR). CIVLAN efficiently localizes pertinent temporal moment in untrimmed videos or subtitles by responding to natural language queries. It addresses three key limitations of previous methods: the extensive need for fully supervised training with labor-intensive annotations, underutilization of auxiliary modalities like subtitles and audio, and the neglection of character identification in dialogues. CIVLAN comprises two main components: the Modality Alignment Network (MAN), which aligns queries to video/subtitle content through contrastive loss minimization, and the Character Identifying Network (CIN), which associates on-screen characters with their references in subtitles and queries. This method eliminates the need for ground-truth temporal moment, significantly reducing manual annotation efforts. Our experiments demonstrate that CIVLAN surpasses previous state-of-the-art methods on the TVR benchmark dataset under weakly-supervised conditions. The paper also includes comprehensive ablation studies and qualitative analyses to showcase the effectiveness of CIVLAN’s components.