<p>Speech is a vital means of communication for humans, through which thoughts, feelings, and ideas are conveyed. However, for some individuals, this fundamental aspect of communication can be a challenge due to a disorder known as stuttering or stammering. This poses challenges for both the speaker and listener. This condition affects millions of people worldwide and can significantly impact their daily lives. Thus, extensive research has been conducted in the realm of speech signal processing to devise efficient methods for identifying and categorizing patterns of stuttered speech. This study focuses on using deep learning techniques, such as CNN and Bi-LSTM models, to classify eight different types of stuttering: fluent, interjections, broken words, prolongations, word repetitions, phrase repetitions, part-word repetitions, and sound repetitions. Accurate classification relies heavily on the feature extraction methods used. In this experiment, we incorporate weighted and max-valued features derived from Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), and MEL-spectrogram features to improve stuttering speech recognition. The primary objective of this paper is to provide a comprehensive analysis of the CNN model with Bi-LSTM, based on the various types of features mentioned earlier. After conducting a thorough analysis, it was observed that using Weighted MFCC (WMFCC) as input features in the CNN model enabled the classification of eight different types of stuttering with notable accuracy, achieving the highest accuracy rate of 91% after 400 epochs. These findings demonstrate the effectiveness and potential of WMFCC features in combination with CNN models for accurately detecting and classifying stuttering patterns in speech. Overall, this research work contributes to a deeper understanding and advancement in the area of speech recognition and its application in classifying different types of stuttering behaviors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep analysis of weighted features utilizing MFCC, MEL-spectrogram, and LPC within the framework of convolution neural network and bidirectional LSTM for automatic stuttering classification

  • Nilanjan Banerjee,
  • Nilambar Sethi,
  • Samarjeet Borah

摘要

Speech is a vital means of communication for humans, through which thoughts, feelings, and ideas are conveyed. However, for some individuals, this fundamental aspect of communication can be a challenge due to a disorder known as stuttering or stammering. This poses challenges for both the speaker and listener. This condition affects millions of people worldwide and can significantly impact their daily lives. Thus, extensive research has been conducted in the realm of speech signal processing to devise efficient methods for identifying and categorizing patterns of stuttered speech. This study focuses on using deep learning techniques, such as CNN and Bi-LSTM models, to classify eight different types of stuttering: fluent, interjections, broken words, prolongations, word repetitions, phrase repetitions, part-word repetitions, and sound repetitions. Accurate classification relies heavily on the feature extraction methods used. In this experiment, we incorporate weighted and max-valued features derived from Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), and MEL-spectrogram features to improve stuttering speech recognition. The primary objective of this paper is to provide a comprehensive analysis of the CNN model with Bi-LSTM, based on the various types of features mentioned earlier. After conducting a thorough analysis, it was observed that using Weighted MFCC (WMFCC) as input features in the CNN model enabled the classification of eight different types of stuttering with notable accuracy, achieving the highest accuracy rate of 91% after 400 epochs. These findings demonstrate the effectiveness and potential of WMFCC features in combination with CNN models for accurately detecting and classifying stuttering patterns in speech. Overall, this research work contributes to a deeper understanding and advancement in the area of speech recognition and its application in classifying different types of stuttering behaviors.