MMT: Transformer for Multi-modal Multi-label Self-supervised Learning
摘要
We proposed a novel network, called Transformer, for Multi-modal Multi-label Self-Supervised Learning (MMT) in the context of video classification. Our approach tackles the input challenges arising from different modalities by incorporating auxiliary tasks that exploit the natural synchronization and correlation between video, audio, and text. To bridge the semantic gap existing among modalities, we perform alignment in both high-dimensional and low-dimensional feature spaces for video-audio and video-text modalities, respectively. Moreover, we introduce a contrastive loss to enhance the proximity of semantically related samples, effectively catering to the specific requirements of multi-modal learning. The experimental results demonstrate that our method achieves competitive performance in various video classification tasks. This demonstrates the effectiveness of our proposed Transformer network for MMT, showcasing its potential in advancing multi-modal learning research.