RE-VC: Robust Zero-Shot Voice Conversion Model for Realistic Environments
摘要
Voice conversion (VC) technology, which enables the transformation of one speaker’s voice to that of another while preserving linguistic content, has seen significant advancements in recent years. However, existing VC models face two major challenges: reduced voice quality when processing noisy speech and diminished voice similarity in zero-shot inference scenarios. In this paper, we propose a novel approach to address these issues by integrating a content encoder with WavLM model to effectively extract linguistic features and an attention-based speaker encoder to capture speaker-specific characteristics. This design enhances both the quality and similarity of the converted speech. Additionally, we introduce a noise augmentation strategy aimed at improving the model’s robustness during training, enabling it to better handle noisy speech in realistic environments. Extensive experiments, including both objective and subjective evaluations, demonstrate that our model surpasses existing methods, offering improved performance and robustness.