Unsupervised Multi-level Search and Correspondence for Generic Voice-Face Feature Spaces
摘要
This paper focuses on the challenge of searching a shared feature space for face and voice modalities. Studies of human intelligence have shown that people can link faces and voices. However, computational intelligence research has paid less attention to the relationship between voice and face. The objective of this work is to identify generic features that are generalized to both visual and audio modalities, enabling the matching of faces and voices. To achieve a better approximation of the feature spaces of two modalities, a multi-level alignment approach is applied to their features. For individual sample pairs, a contrast learning approach is exploited. For the overall feature distribution, an optimal transport method is used. The impact of pre-trained weights on this task is explored. Compared to the simple contrastive learning method, inclusion of overall feature alignment improves both verification and matching accuracy. Extensive experimental results suggest that the overall distribution alignment is useful for the cross-model feature matching task. Also, the use of pre-trained parameters can improve the results in certain situations.