Sequential Consistency Matters: Boosting Video Sequence Verification with Teacher Multimodal Transformer
摘要
Video sequence verification (SV), which aims to determine whether the procedures within two videos are consistent at the step level, has received significant attention in recent years. This paper introduces a novel approach to boost the SV task by leveraging multimodal data, including video frames and accompanying textual narrations. The core of the proposed method is the Teacher Multimodal Transformer, designed to facilitate the learning of matching information between different modalities and improve the SV task's performance through a pseudo-label generation algorithm. Additionally, a cross-grained contrastive learning loss is introduced to effectively capture relevant information between coarse-grained and fine-grained features. The paper demonstrates the efficacy of the proposed method on three widely used SV datasets, including CSV, Diving-SV, and COIN-SV. The proposed method has 2.23%, 1.56%, 2.6% improved to the conventional methods on these benchmarks under weakly supervised setting.