Character Identifying Video Language Alignment Network for Weakly-Supervised Video-Subtitle Moment Retrieval
摘要
This paper introduces the Character Identifying Video Language Alignment Network (CIVLAN), a labor-efficient approach for weakly-supervised Video-Subtitle Moment Retrieval (VSMR). CIVLAN efficiently localizes pertinent temporal moment in untrimmed videos or subtitles by responding to natural language queries. It addresses three key limitations of previous methods: the extensive need for fully supervised training with labor-intensive annotations, underutilization of auxiliary modalities like subtitles and audio, and the neglection of character identification in dialogues. CIVLAN comprises two main components: the Modality Alignment Network (MAN), which aligns queries to video/subtitle content through contrastive loss minimization, and the Character Identifying Network (CIN), which associates on-screen characters with their references in subtitles and queries. This method eliminates the need for ground-truth temporal moment, significantly reducing manual annotation efforts. Our experiments demonstrate that CIVLAN surpasses previous state-of-the-art methods on the TVR benchmark dataset under weakly-supervised conditions. The paper also includes comprehensive ablation studies and qualitative analyses to showcase the effectiveness of CIVLAN’s components.