Video action recognition plays a crucial role in numerous applications. However, current research encounters significant challenges, such as the difficulty in extracting features that maintain long-term connections within the video feature space and the inefficiency of attention computation. Therefore, we propose an innovative feature fusion method based on attention mechanism. This method employs the Transformer to create a Key Frame and Key Patch Selection module alongside a Small-Big Patch Transformer module. These components efficiently establish relationships between features with long-term connections in video data. By integrating these modules within the Transformer and combining them with Convolutional Neural Networks, our approach leverages the strengths of different frameworks. This integration significantly enhances the computational efficiency and accuracy. The experiments conducted on the generic video datasets demonstrate that our model surpasses the performance of most previous efforts at the same input scale.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FPSFT: Frame-Patch-Select Fusion Transformer for Action Recognition

  • Yilong Xiao,
  • Xin Jing,
  • Shiyu Yu,
  • Jingyu Liu,
  • Keming Mao,
  • Xiaochun Yang,
  • Bin Zhang

摘要

Video action recognition plays a crucial role in numerous applications. However, current research encounters significant challenges, such as the difficulty in extracting features that maintain long-term connections within the video feature space and the inefficiency of attention computation. Therefore, we propose an innovative feature fusion method based on attention mechanism. This method employs the Transformer to create a Key Frame and Key Patch Selection module alongside a Small-Big Patch Transformer module. These components efficiently establish relationships between features with long-term connections in video data. By integrating these modules within the Transformer and combining them with Convolutional Neural Networks, our approach leverages the strengths of different frameworks. This integration significantly enhances the computational efficiency and accuracy. The experiments conducted on the generic video datasets demonstrate that our model surpasses the performance of most previous efforts at the same input scale.