错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SPViM: Sparse Pyramid Video Representation Learning Framework for Fine-Grained Action Retrieval

  • Lutong Wang,
  • Chenglei Yang,
  • Hongqiu Luan,
  • Wei Gai,
  • Wenxiu Geng,
  • Yawen Zheng

摘要

Existing research has achieved remarkable success for video-based action understanding. However, current researches mainly focus on recognizing external actions at coarse-grained, with less attention paid to the fine-grained action understanding, thus impeding the precise localization and retrieval of internal content. To this end, we propose a Sparse Pyramid Video representation learning framework (SPViM), aiming to achieve frame-to-frame retrieval related to high-level action semantics. Firstly, an appearance encoder is introduced to construct independent visual descriptors for each input frame, where a shift window mechanism captures the underlying inter-frame nuances. Secondly, a temporal encoder containing the sparse self-attention and multi-granularity local context awareness mechanism were constructed to comprehensively describe the action hierarchy. Herein, inspired by the human brain cognitive process when retrieving specific content, we design a set of sparse constraints to guide self-attention gradually converge from global sparse to local dense centered on the target frame. Furthermore, we develop a Transformer-based temporal pyramid structure to integrate multi-scale spatio-temporal features, thereby generating comprehensive and discriminative video frame representations. Extensive experiments show that our fine-grained video retrieval method with SPViM architecture outperforms the state-of-the-art method on three challenging datasets.