Domain Adaptation for Speaker Verification Based on Self-supervised Learning with Adversarial Training
摘要
Speaker verification models trained on a single domain have difficulty keeping performance on new domain data. Adversarial training maps different domain data to the same subspace to handle this problem. However, adversarial training only uses domain labels on the target domain and does not mine its speaker information. To improve the domain adaptation performance for speaker verification, we propose a joint training strategy for adversarial training and self-supervised learning. In our method, adversarial training adapts knowledge from the source domain to the target domain, while self-supervised learning obtains speech representations from unlabeled utterances. Further, our self-supervised learning only uses positive pairs to avoid false negative samples. The proposed joint training strategy enables adversarial training to guide self-supervised learning to focus on speaker verification tasks. Experiments show our proposed method outperforms other domain adaptation methods.