Temporal text localization (TTL) task refers to identify a segment within a long untrimmed video that semantically matches a given textual query. However, most methods require extensive manual annotation of temporal boundaries for each query, which restricts their scalability and practicality in real-world applications. Moreover, modeling temporal context information is particularly crucial for TTL task. In this paper, a Vision Token Rolling Transformer for weakly supervised temporal text localization (VTR-former) is developed. VTR-former does not rely on predefined temporal boundaries during training or testing. It significantly improves the performance of the model in temporal information capture and feature representation by rolling vision tokens and utilizing advanced feature learning modules based on the transformer. Experiments on two challenging benchmarks, including Charades-STA and ActivityNet Captions, demonstrate that VTR-former outperforms the baseline network and achieves the leading performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VTR-Former: Vision Token Rolling Transformer for Weakly Supervised Temporal Text Localization

  • Zeyu Xi,
  • Xinlang Zhou,
  • Zilin Liu,
  • Lifang Wu

摘要

Temporal text localization (TTL) task refers to identify a segment within a long untrimmed video that semantically matches a given textual query. However, most methods require extensive manual annotation of temporal boundaries for each query, which restricts their scalability and practicality in real-world applications. Moreover, modeling temporal context information is particularly crucial for TTL task. In this paper, a Vision Token Rolling Transformer for weakly supervised temporal text localization (VTR-former) is developed. VTR-former does not rely on predefined temporal boundaries during training or testing. It significantly improves the performance of the model in temporal information capture and feature representation by rolling vision tokens and utilizing advanced feature learning modules based on the transformer. Experiments on two challenging benchmarks, including Charades-STA and ActivityNet Captions, demonstrate that VTR-former outperforms the baseline network and achieves the leading performance.