Enhancing Domain Adaptation in Speaker Verification via Partially Shared Adversarial Network
摘要
Speaker recognition technology has achieved excellent performance on public datasets, with speaker embeddings like x-vectors demonstrating remarkable effectiveness in representing speaker characteristics. However, most existing research is limited to training and testing within the same environment, leading to domain mismatches in complex scenarios. To address this, we propose a domain-adaptive speaker verification framework called the Partially Shared Adversarial Network (PSAN). This framework uses a shared representation and an adaptive loss weighting mechanism to learn both speaker-general and domain-specific knowledge from multiple target domains, enhancing the discriminability of speaker embeddings. In our experiments, LibriSpeech served as the source domain, and target domain datasets were created by adding noise. Results show that PSAN not only excels in the source domain but also demonstrates superior domain adaptation capabilities on the target domain compared to state-of-the-art models.