SPLGAN-TTS: Learning Semantic and Prosody to Enhance the Text-to-Speech Quality of Lightweight GAN Models
摘要
Autoregressive-based models have proven effective in speech synthesis; however, numerous parameters and slow inference limit their applicabili ty. Though non-autoregressive models can resolve these issues, speech synthesis quality is unsatisfactory. This study employed a tree-based structure to enhance the learning of semantic and prosody information using a lightweight model. A Variational Encoder (VAE) is used for the generator architecture, and a novel normalizing-flow module is used to enhance the complexity of the VAE-generated distribution. We also developed a speech discriminator with a multi-length architecture to reduce computational overhead as well as multiple auxiliary losses to assist in model training. The proposed model is smaller than existing state-of-the-art models, and synthesis performance is faster, particularly when applied to longer texts. Despite the fact that the proposed model is roughly 30% smaller than FastSpeech2 [1], its mean opinion score surpasses FastSpeech2 as well as other models.