错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VANet: A New Network for Multi-modal Self-supervised Learning from Video and Audio

  • Xingrui Liu,
  • Chen Zhang,
  • Zeming Feng,
  • Jiwen Dong,
  • Sijie Niu,
  • Xizhan Gao

摘要

A new Network for Multi-modal Self-Supervised Learning from Video and Audio (VANet) is proposed for video action recognition. To overcome the input issues of different modalities, we establish auxiliary tasks by exploiting the natural synchronization and correlation between video and audio, and align them at both the segment-level representation and the video/audio-level representation. In this work, we introduce a contrastive loss for joint embedding learning of the two input modalities, which is improved to enhance the proximity of semantically relevant samples and meet the special needs of multi-modal learning. In addition, we identify and remove false negative samples from the negative set. Our method achieves competitive results in video action recognition tasks and demonstrates that joint multi-modal learning is superior to single modality.