Video sequence verification (SV), which aims to determine whether the procedures within two videos are consistent at the step level, has received significant attention in recent years. This paper introduces a novel approach to boost the SV task by leveraging multimodal data, including video frames and accompanying textual narrations. The core of the proposed method is the Teacher Multimodal Transformer, designed to facilitate the learning of matching information between different modalities and improve the SV task's performance through a pseudo-label generation algorithm. Additionally, a cross-grained contrastive learning loss is introduced to effectively capture relevant information between coarse-grained and fine-grained features. The paper demonstrates the efficacy of the proposed method on three widely used SV datasets, including CSV, Diving-SV, and COIN-SV. The proposed method has 2.23%, 1.56%, 2.6% improved to the conventional methods on these benchmarks under weakly supervised setting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sequential Consistency Matters: Boosting Video Sequence Verification with Teacher Multimodal Transformer

  • Yaning Zhao,
  • Xun Jiang,
  • Jingran Zhang,
  • Guofeng Yi

摘要

Video sequence verification (SV), which aims to determine whether the procedures within two videos are consistent at the step level, has received significant attention in recent years. This paper introduces a novel approach to boost the SV task by leveraging multimodal data, including video frames and accompanying textual narrations. The core of the proposed method is the Teacher Multimodal Transformer, designed to facilitate the learning of matching information between different modalities and improve the SV task's performance through a pseudo-label generation algorithm. Additionally, a cross-grained contrastive learning loss is introduced to effectively capture relevant information between coarse-grained and fine-grained features. The paper demonstrates the efficacy of the proposed method on three widely used SV datasets, including CSV, Diving-SV, and COIN-SV. The proposed method has 2.23%, 1.56%, 2.6% improved to the conventional methods on these benchmarks under weakly supervised setting.