错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-scale Multi-modal Multi-dimension Joint Transformer for Two-Stream Action Classification

  • Lin Wang,
  • Ammar Hawbani,
  • Yan Xiong

摘要

Multi-modal, multi-scale, and multi-dimensional (spatiotemporal) video representation learning have each been studied adequately in its own form, respectively, but rather in an isolated way from each other, not yet jointly. It is well known in statistical machine learning that joint data distributions provide new information that cannot be achieved by its individual components. Therefore we propose M \(^3\) T : a Multi-scale Multi-modal Multi-dimension (M \(^3\) ) joint Transformer model for two-stream video representation learning, which is built upon a two-stream multi-scale vision transformer backbone. M \(^3\) T is densely augmented with three attention modules, which are mutually orthogonal against each other, at each down-sampling layer of the backbone. Experiments conducted on the Kinetics 400 data set demonstrate the effectiveness of the proposed method. The qualitative performance also demonstrates that our model can learn more informative complementary representation.