错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning’s Impact on Speech Synthesis for Mobile Devices

  • Ronit Hemant Chougule,
  • Akshay Kamkhalia

摘要

In recent years, speech synthesis has become an integral part-of-speech-based human–machine interfaces on mobile devices, sparking significant research interest. While traditional approaches like formant-based synthesis, unit-selection synthesis, and Hidden Markov Model (HMM)-based synthesis have been widely used, HMM-based synthesis, a type of Statistical Parametric Speech Synthesis (SPSS), has gained popularity due to its compact model size and flexibility in adjusting voice characteristics. However, the synthetic speech produced by HMM-based synthesis often sounds muffled due to over-smoothing, indicating the need for improvement. Implementing speech synthesis on mobile devices also presents challenges, such as limited memory resources, processing capacity, and real-time responsiveness. To tackle these challenges, researchers have investigated the use of deep learning models to improve voice quality, enhance prediction performance, and optimize embedded implementation without compromising voice quality in HMM-based synthesis systems. Deep learning models have also been applied in different stages of conventional SPSS to significantly reduce the memory footprint. These findings suggest that deep learning has a significant impact on speech synthesis for mobile devices, with the potential for further improvements in voice quality and efficient embedded implementation. Thus, this area presents a promising research direction for the future.