错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MMT: Transformer for Multi-modal Multi-label Self-supervised Learning

  • Jiahe Wang,
  • Jia Li,
  • Xingrui Liu,
  • Xizhan Gao,
  • Sijie Niu,
  • Jiwen Dong

摘要

We proposed a novel network, called Transformer, for Multi-modal Multi-label Self-Supervised Learning (MMT) in the context of video classification. Our approach tackles the input challenges arising from different modalities by incorporating auxiliary tasks that exploit the natural synchronization and correlation between video, audio, and text. To bridge the semantic gap existing among modalities, we perform alignment in both high-dimensional and low-dimensional feature spaces for video-audio and video-text modalities, respectively. Moreover, we introduce a contrastive loss to enhance the proximity of semantically related samples, effectively catering to the specific requirements of multi-modal learning. The experimental results demonstrate that our method achieves competitive performance in various video classification tasks. This demonstrates the effectiveness of our proposed Transformer network for MMT, showcasing its potential in advancing multi-modal learning research.