A Novel and Intelligent Approach for Indian Locale Based Text-to-Speech Model by Hybridizing Wave Net and Wave Glow with Mel-Spectrogram Analysis
摘要
This paper introduces a innovative architecture for text-to-speech synthesis, comprising a recurrent speaker encoder, a sequence-to-sequence synthesizer, and various vocoder options like Griffin-Lim, WaveNet, and WaveGlow. The speaker encoder extracts a fixed-dimensional vector from speech signals, leveraging a text-independent speaker verification model trained on reference speech samples from the target speaker, resulting in the generation of a speaker embedding vector. The training data includes text transcripts paired with target audio, and transfer learning is facilitated through a pre-trained speaker encoder. The vocoder component plays a critical role in converting synthesized mel-spectrograms into time-domain waveforms, exploring various vocoder options such as Griffin-Lim, WaveNet, and WaveGlow. This paper aims to evaluate the performance of two vocoder variants, including the hybridization of Waveglow and Wavenet. The research seeks to advance the field of multispeaker text-to-speech synthesis by enhancing the quality of generated speech and enabling speaker adaptation. The paper explores various vocoder options, including Griffin-Lim, which iteratively estimates missing phase information, while WaveNet employs up-sampling layers to achieve high-resolution audio. WaveGlow efficiently synthesizes high-quality audio without the need for autoregression by combining insights from Glow and WaveNet. The research that we have performed focuses on the accent-based speech within the context of a novel text-to-speech synthesis. These findings hold significant promise for applications in speech synthesis, voice cloning, and natural language processing, contributing to the development of more versatile and high-quality text-to-speech systems.