One-Shot Voice Conversion Based on Style Generative Adversarial Networks with ESR and DSNet
摘要
This paper proposes a novel one-shot voice conversion (VC) method called DS-ESR-StyleGAN-VC, which encompasses several innovations to address the challenges faced by StarGAN-VC. Firstly, we adopt ESR network in the generator to extract deep features, effectively solving the problem of semantic content corruption in StarGAN-VC. Secondly, we leverage the advantages of the dense weighted normalized shortcut employed by DSNet, which circumvents the performance degradation and gradient disappearance caused by the increasing convolutional layers. The DSNet network is integrated into the middle of the encoder and decoder of generator to further extract the spliced features, further enhancing VC quality. Thirdly, we remove the classifier module in StarGAN-VC and use a style encoder to extract speaker style features in order to improve speaker similarity. Moreover, our proposed method can naturally support one-shot VC. Experiments show that our proposed method consistently outperforms the competitive StarGAN-VC in terms of semantic content completeness, naturalness and speaker similarity under the many-to-many setting; our proposed method is superior to the competitive StarGAN-ZSVC in terms of naturalness under the one-shot setting.