错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatiotemporal Representation Enhanced ViT for Video Recognition

  • Min Li,
  • Fengfa Li,
  • Bo Meng,
  • Ruwen Bai,
  • Junxing Ren,
  • Zihao Huang,
  • Chenghua Gao

摘要

Vision Transformers (ViTs) are promising for solving video-related tasks, but often suffer from computational bottlenecks or insufficient temporal information. Recent advances in large-scale pre-training show great potential for high-quality video representation, providing new remedies to transformer limitations. Inspired by this, we propose a SpatioTemporal Representation Enhanced Vision Transformer (STRE-ViT), which follows a two-stream paradigm to fuse large-scale pre-training visual prior knowledge and video-level temporal biases in a simple and effective manner. Specifically, one stream employs a well-pretrained ViT with rich vision priors to alleviate data requirements and learning workload. Another stream is our designed spatiotemporal interaction stream, which first models video-level temporal dynamics and then extracts fine-grained and salient spatiotemporal representations by introducing appropriate temporal bias. Through this interaction stream, the model capacity of ViT is enhanced for video spatiotemporal representations. Moreover, we provide a fresh perspective to adapt well-pretrained ViT for video recognition. Experimental results show that STRE-ViT learns high-quality video representations and achieves competitive performance on two popular video benchmarks.