An Efficient Speech Synthesizer: A Hybrid Monotonic Architecture for Text-to-speech via VAE & LPC-Net with Independent Sentence Length
摘要
This research delineates the development of a hybrid architecture for text-to-speech (TTS) synthesis, termed the Efficient Speech Synthesizer (ESS). ESS leverages an end-to-end training paradigm to holistically optimize all the parameters, thereby facilitating the synthesis of high-fidelity speech with remarkable efficiency, irrespective of sentence length. The architecture integrates state-of-the-art feed-forward methodologies, including variational autoencoders and linear predictive coding techniques, which impose monotonic constraints on sequence alignment while incurring negligible computational overhead. The primary objective of this study is to engineer a TTS model characterized by minimal complexity and latency, rendering it suitable for deployment across a broad spectrum of computational platforms, ranging from high-performance systems to resource-constrained IoT devices and smartphones.