The research focuses on analyzing intonation and acoustic features of speech using machine learning techniques. The relevance of the research is due to the attempt to search for a methodology to create a system based on the combination of prosodic and spectral features and the selection of an optimal classifier to identify whether there is an accent in speech. The authors discuss the advantages and constraints of the most commonly used machine learning algorithms in binary classification tasks: the K-nearest neighbor (KNN) algorithm, support vector method (SVM), decision tree algorithms (DTC), and machine learning logistic regression algorithms that were used during the research. As a result of analyzing the importance of features, the authors determined the “basic combination of features,” which demonstrates the highest accuracy. The scientific novelty of this research is determined by using machine learning techniques to create a system based on a combination of prosodic and spectral features and select the optimal classifier for identifying the presence and absence of an accent in speech.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Algorithms in the Speech Analysis: Intonation and Acoustic Characteristics Comparative Study

  • Oxana V. Goncharova,
  • Svetlana A. Khaleeva,
  • Lola A. Kaufova,
  • Almira M. Kazieva,
  • Natalia V. Khomovich

摘要

The research focuses on analyzing intonation and acoustic features of speech using machine learning techniques. The relevance of the research is due to the attempt to search for a methodology to create a system based on the combination of prosodic and spectral features and the selection of an optimal classifier to identify whether there is an accent in speech. The authors discuss the advantages and constraints of the most commonly used machine learning algorithms in binary classification tasks: the K-nearest neighbor (KNN) algorithm, support vector method (SVM), decision tree algorithms (DTC), and machine learning logistic regression algorithms that were used during the research. As a result of analyzing the importance of features, the authors determined the “basic combination of features,” which demonstrates the highest accuracy. The scientific novelty of this research is determined by using machine learning techniques to create a system based on a combination of prosodic and spectral features and select the optimal classifier for identifying the presence and absence of an accent in speech.