An Unsupervised Domain Adaptation Method Based on Distribution Alignment for Speaker Verification
摘要
Existing deep embedding learning based speaker verification (SV) methods may suffer from domain mismatch issue caused by different languages, speaking styles, and environmental noises, leading to poor performance in unseen domains. Moreover, it is generally impractical to collect the utterances of the same speaker from different domains. To address this issue, we propose a distribution alignment based unsupervised domain adaptation method for SV under the multi-task learning framework. Specifically, the embeddings of the labeled source domain are learned through a supervised learning task using cross-entropy loss. The embeddings of the unlabeled target domain are learned via a contrastive method based on swapped prediction, where online clustering is used to produce assignments for different views of the same utterance. By enforcing invariance between assignments, an effective self-supervised learning task is performed for the target domain, with the cluster assignments serving as the pseudo-classes. Under the multi-task learning framework, a distribution alignment loss is further proposed for UDA, which introduces the consistency of intra- and inter-class covariance between source domain classes and target domain pseudo-classes. Extensive experiments on benchmark NIST SRE16 and SRE18 datasets demonstrate that effectiveness of the proposed method.