错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of Machine Learning Methods for Speech Disfluencies’ Classification

  • Nitin Mohan Sharma,
  • Prasant Kumar Mahapatra,
  • Vaibhav Gandhi

摘要

Speech is essential for communication as it allows us to express ourselves and enables us to use the systems that are speech-based. Disfluency is referred to as any interruption in speaking and can often adversely impact an individual's life quality. The paper presents an experimental study of several methods for identifying and categorizing the speech disfluencies. More specifically, this study discusses two disfluencies: prolongation and repetition. We have investigated various machine learning algorithms using the University College London Archive of Stuttered Speech dataset, a popular disfluency dataset created by University College London. Manual segmentation, although a time-consuming approach, has been performed on ten speech files from the dataset, generating a total of 335 disfluent speech samples to train the classifiers. Linear Predictive Cepstral Coefficients and Mel-Frequency Cepstral Coefficients (MLCCs) are two feature extraction methods that have been applied. A comparison of many classifiers and their variations reveals that subspace kNN achieves the highest test accuracy of 87.1% with MFCC features. The future plan is to develop a system for automatically segmenting disfluent speech and classify it.