错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

One-Shot Voice Conversion Based on Style Generative Adversarial Networks with ESR and DSNet

  • Yanping Li,
  • Lei Pan,
  • Xiangtian Qiu,
  • Zeyu Yang,
  • Zhicheng Tan,
  • Bo Qian

摘要

This paper proposes a novel one-shot voice conversion (VC) method called DS-ESR-StyleGAN-VC, which encompasses several innovations to address the challenges faced by StarGAN-VC. Firstly, we adopt ESR network in the generator to extract deep features, effectively solving the problem of semantic content corruption in StarGAN-VC. Secondly, we leverage the advantages of the dense weighted normalized shortcut employed by DSNet, which circumvents the performance degradation and gradient disappearance caused by the increasing convolutional layers. The DSNet network is integrated into the middle of the encoder and decoder of generator to further extract the spliced features, further enhancing VC quality. Thirdly, we remove the classifier module in StarGAN-VC and use a style encoder to extract speaker style features in order to improve speaker similarity. Moreover, our proposed method can naturally support one-shot VC. Experiments show that our proposed method consistently outperforms the competitive StarGAN-VC in terms of semantic content completeness, naturalness and speaker similarity under the many-to-many setting; our proposed method is superior to the competitive StarGAN-ZSVC in terms of naturalness under the one-shot setting.