Advancements in Grapheme-to-Phoneme Conversion Models for Speech Synthesis
摘要
Grapheme-to-phoneme (G2P) conversion, which maps written symbols (graphemes) to units of sound (phonemes), is a fundamental component in text-to-speech (TTS) synthesis, automatic speech recognition (ASR), and language learning systems, transforming written text into phonetic representations. This paper reviews the evolution of G2P methodologies, from early list look-up and rule-based systems to statistical approaches and modern neural network architectures, including recurrent networks, attention mechanisms, and transformers. Recent advancements in G2P research are discussed, with a focus on emerging trends such as sentence-level modeling, integration of acoustic data, and multilingual frameworks. Key challenges, including handling out-of-vocabulary words, homograph disambiguation, and speaker variability, are analyzed, along with potential solutions leveraging pre-trained models and self-supervised learning. By highlighting these developments, this paper aims to inform the design of resource-efficient, adaptable G2P systems capable of addressing the complexities of diverse real-world applications.