Dialect Classification for Hindi and Bhojpuri Languages
摘要
Voice interfaces are gaining importance with the advent of speech technology. Automatic speech recognition (ASR) algorithms and strategies are widely employed for applications such as banking, agriculture, and dialog management systems. In a multilingual country like India, where language is diverse, conversational speech plays a crucial role in the system performance of ASR. Hence, identification and classification of dialects are intuitively expected to improve the system performance of an ASR. The article describes an end-to-end system to identify and classify the dialects of Hindi and Bhojpuri languages using the information present in speech signals. Bhojpuri is a language that is widely spoken in Bihar, India. Bhojpuri and Hindi share similarities in grammar and pronunciation. The focus of the work is to explore the possibility of a Hindi and Bhojpuri dialect identifier (DID) that is based on the intonational differences between the dialects. Intonation patterns will be explored to find cues or clues to identify the dialects for the Hindi and Bhojpuri languages.