Text-to-Speech Conversion Using Concatenative Approach for Gujarati Language
摘要
Speech is often regarded as the primary and innate mode of communication used by individuals within the human species. For the last three decades, there has been a concerted effort by researchers to develop computers capable of comprehending and engaging in human-like conversation. In contrast to English and other European languages such as French and Spanish, there is a notable dearth of study conducted in Indian languages. Speech serves as a means of communicating information, as well as expressing emotions and sentiments. Voice synthesis refers to the methodology used in transforming a particular input text into synthetic voice. Typically, the text-to-speech (TTS) system has two distinct stages. One of the primary techniques used in this study is text analysis, which involves converting the given input text into a phonetic or other linguistic representation. The second aspect pertains to the production of speech waveforms. Speech synthesis, also known as text-to-speech synthesis, refers to the process of artificially generating human-like voice from written text, specifically in the context of Gujarati language. This paper presents an ongoing study in the field of text-to-speech synthesis, focusing primarily on the concatenative approach.