Survey Paper on Text Echo: Personalized TTS System
摘要
The proposed system introduces a text-to-speech (TTS) solution that converts written content into audio. This technology can process text from diverse sources such as PDFs and images to generate natural-sounding spoken words. This approach utilizes OCR technology to extract text with precision and implements advanced deep learning models, including Tacotron 2 and WaveNet, to generate the speech output. A key feature of this invention is the ability of users to record their own voices, enabling the generation of personalized speech that mimics their unique vocal characteristics and emotional tones. Designed with an intuitive user interface, TTS supports multiple languages and offers real-time processing. This innovative solution significantly enhances accessibility for individuals with disabilities, language learners, and reading difficulties, thereby providing versatile and engaging auditory experiences.