Robust Contrastive Learning Against Audio-Visual Noisy Correspondence
摘要
Recent efforts have focused on training audio-visual pairs through self-supervised contrastive learning, which relies on the assumption of audio-visual correspondence (AVC). This assumption posits that positive pairs consist of audio and visual from the same video, while negative pairs are formed from different videos. However, this assumption is too strict and may be unreliable in practice. This unreliable assumption inevitably introduces two types of noisy correspondence. False positive pairs arise from weak AVC caused by invisible-sounding objects or background noise. Conversely, false negative pairs arise from strong AVC caused by random pairing. In this paper, we focus on the visual sound localization task, aiming to localize the visual regions that emit sound. To address the issue of noisy correspondence in visual sound localization, an optimized soft contrastive loss is proposed to alleviate the impact of false positives. Additionally, the hard contrastive set mixing strategy is utilized to suppress the effect of false negatives. Experimental results demonstrate that our methods significantly reduce the impact of noisy correspondences and achieve competitive results on standard benchmarks. Furthermore, the proposed method shows potential for generalization to other multi-modal tasks based on contrastive learning.