A Multi-scale Multi-modal Multi-dimension Joint Transformer for Two-Stream Action Classification
摘要
Multi-modal, multi-scale, and multi-dimensional (spatiotemporal) video representation learning have each been studied adequately in its own form, respectively, but rather in an isolated way from each other, not yet jointly. It is well known in statistical machine learning that joint data distributions provide new information that cannot be achieved by its individual components. Therefore we propose M \(^3\) T : a Multi-scale Multi-modal Multi-dimension (M \(^3\) ) joint Transformer model for two-stream video representation learning, which is built upon a two-stream multi-scale vision transformer backbone. M \(^3\) T is densely augmented with three attention modules, which are mutually orthogonal against each other, at each down-sampling layer of the backbone. Experiments conducted on the Kinetics 400 data set demonstrate the effectiveness of the proposed method. The qualitative performance also demonstrates that our model can learn more informative complementary representation.