An Efficient Momentum Framework for Face-Voice Association Learning
摘要
Cross-modal face-voice association is an active field that utilizes biometric features for cross-modal information retrieval. The primary approach for addressing this task involves utilizing contrastive learning to construct a modality-agnostic subspace. However, many existing contrastive learning methods in cross-modal research tend to neglect the significance of symmetrical information within heterogeneous data. This oversight leads to the generation of different negative examples for each identity in a random mini-batch. Furthermore, the length of negative examples in contrastive learning is coupled with the mini-batch size and is limited by the GPU memory size. To address these issues, this paper introduces an innovative Cross-Modal Momentum Contrast (CMMC) algorithm, which leverages queues to provide sufficient and symmetric information. Moreover, we propose an update strategy to maintain the consistency of negative example information throughout the training process. By combining the operations mentioned above, our proposed CMMC can effectively improve the correlation between face and voice data. Extensive experiments conducted on two datasets confirm the superiority of our framework and demonstrate its competitive performance compared to state-of-the-art methods.