TStarGANv2-VC: Non-parallel Multi-domain Transformer Based StarGANv2 Voice Conversion
摘要
Voice Conversion (VC) involves altering the speaking style of a voice while maintaining the original linguistic content. Researchers commonly utilize advanced deep generative frameworks for VC applications, with Generative Adversarial Networks (GANs) being a favored choice. Methods such as CycleGAN-VC and StarGAN-VC focus on producing speech that closely resembles the target speaker’s voice. Although these techniques are effective, they can face challenges such as inconsistencies and artifacts in the modified speech. To tackle these issues, StarGANv2-VC was created as a system for unsupervised, non-matching, multi-to-multi speech transformation. It leverages adversarial loss based on the source classifier and perceptual loss to achieve varied and high-quality speech style adaptations. However, challenges remain in handling long-range dependencies and global structures in voice signals for maintaining naturalness and quality. Therefore, this article introduces a new Transformer-based StarGANv2-VC model (TStarGANv2-VC), which aims to disentangle variables specific to a particular domain or manner in VC. This new method utilizes query weights and a new learning method for extrication, along with a Domain-Representative Input (DRI) for improving the conversion process. A domain-specificity extrication criterion is also defined to assess the disentanglement quality. Performance analysis on various datasets shows the efficiency of TStarGANv2-VC compared to existing VC models. Test results on a non-matching, multi-to-multi speech transformation task show that TStarGANv2-VC generates lifelike speech that closely resembles the performance of leading VC models.