VANet: A New Network for Multi-modal Self-supervised Learning from Video and Audio
摘要
A new Network for Multi-modal Self-Supervised Learning from Video and Audio (VANet) is proposed for video action recognition. To overcome the input issues of different modalities, we establish auxiliary tasks by exploiting the natural synchronization and correlation between video and audio, and align them at both the segment-level representation and the video/audio-level representation. In this work, we introduce a contrastive loss for joint embedding learning of the two input modalities, which is improved to enhance the proximity of semantically relevant samples and meet the special needs of multi-modal learning. In addition, we identify and remove false negative samples from the negative set. Our method achieves competitive results in video action recognition tasks and demonstrates that joint multi-modal learning is superior to single modality.