A speaker’s perspective cannot be fully expressed by phoneme segment in a speech, but it also comes along with the intonation. In the Nepali language, declarative questions are uttered with rising intonation and are commonly used in daily conversation. In this paper, training on various machine learning models has been performed to classify declarative sentences and questions. The dataset consists of 700 utterances with ~ 300 declarative statements (falling intonation) and ~ 400 declarative question statements (rising intonation) which are either recorded in the studio or taken from the “Large Nepali ASR training dataset” from openslr.org. The dataset was pre-processed and annotated using PRAAT. Cepstral features MFCCs of every utterance in the data set were calculated and put to use in training five different supervised machine learning models. The data set is split into two parts: 70% for training and 30% for testing purposes, respectively. The comparative study shows that the Random Forest model performed better with the highest evaluation score of ~ 0.87% and the Naïve Bayes model achieved a minimal accuracy of approximately 0.69%. These findings have the potential to enhance the authenticity of Text-to-Speech, Automatic–Speech-Recognition, and Speech-to-Text systems in languages with limited linguistic resources, like Nepali.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Intonation Analysis and Classification of Declarative Questions and Statements in Nepali Language using Machine Learning Approaches

  • Ranjit Subba,
  • Roshan Kumar Prasad,
  • Pratika Rai,
  • Sharad Sinha

摘要

A speaker’s perspective cannot be fully expressed by phoneme segment in a speech, but it also comes along with the intonation. In the Nepali language, declarative questions are uttered with rising intonation and are commonly used in daily conversation. In this paper, training on various machine learning models has been performed to classify declarative sentences and questions. The dataset consists of 700 utterances with ~ 300 declarative statements (falling intonation) and ~ 400 declarative question statements (rising intonation) which are either recorded in the studio or taken from the “Large Nepali ASR training dataset” from openslr.org. The dataset was pre-processed and annotated using PRAAT. Cepstral features MFCCs of every utterance in the data set were calculated and put to use in training five different supervised machine learning models. The data set is split into two parts: 70% for training and 30% for testing purposes, respectively. The comparative study shows that the Random Forest model performed better with the highest evaluation score of ~ 0.87% and the Naïve Bayes model achieved a minimal accuracy of approximately 0.69%. These findings have the potential to enhance the authenticity of Text-to-Speech, Automatic–Speech-Recognition, and Speech-to-Text systems in languages with limited linguistic resources, like Nepali.