Stuttering, a speech impediment that significantly impacts an individual’s personality and daily life, is a critical research area requiring technological solutions. Therefore, bringing our attention, we analyzed speech utterances of people who stutter, using recordings from the publicly available Sep-28k dataset, which consists of natural conversations from podcasts. A major focus has been put on analyzing the fact that data imbalance plays a role in learning the characteristics of a speech utterance. We have used a sampling technique by considering the middle value of samples calculated from the minority and the majority class and this serves as a combination of up-sampling and down-sampling. This helps us to avoid bias in the class-wise characteristic learning of the model. The proposed approach focuses on two feature extraction techniques, specifically Mel-Frequency Cepstral Coefficients (MFCCs), del-del MFCC and a comprehensive analysis using the openSMILE toolkit, selecting only the ComParE2016 features. We conducted a comparative study between various machine learning models like random forest (RF), support vector machine (SVM) and time-delayed neural network (TDNN) using both features. Our findings indicated that the stutter class outperformed the no stutter class by a significant margin for each classifier in binary classification. This prompted us to conduct two experiments: one with all classes and another with only stutter classes. The experimental analysis revealed an average F1-score with del-del MFCC features of 85.38% with RF, 90.10% with SVM when all classes are considered. Furthermore, 93.29% with RF, 93.71% with SVM when only stutter classes are considered. On the contrary, when we use ComParE2016 features we achieve an average accuracy of 86.79% with RF and 87.77% with SVM considering all classes. Whereas, 95.10% with RF, 94.34% with SVM when only stutter classes are considered. When TDNN was used with del-del MFCC, we achieved an accuracy of 79.02% and 92% with ComParE2016, implying better performance because of better feature extraction. It is noteworthy that throughout all the experiments, sound repetition is the best performing class indicating consistency in the learning process. The results achieved through this sampling technique improve the stutter detection and classification by a significant margin.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Does Data Balancing Impact Stutter Detection and Classification?

  • Ashita Batra,
  • Pratyush Shrivastava,
  • Pradip K. Das

摘要

Stuttering, a speech impediment that significantly impacts an individual’s personality and daily life, is a critical research area requiring technological solutions. Therefore, bringing our attention, we analyzed speech utterances of people who stutter, using recordings from the publicly available Sep-28k dataset, which consists of natural conversations from podcasts. A major focus has been put on analyzing the fact that data imbalance plays a role in learning the characteristics of a speech utterance. We have used a sampling technique by considering the middle value of samples calculated from the minority and the majority class and this serves as a combination of up-sampling and down-sampling. This helps us to avoid bias in the class-wise characteristic learning of the model. The proposed approach focuses on two feature extraction techniques, specifically Mel-Frequency Cepstral Coefficients (MFCCs), del-del MFCC and a comprehensive analysis using the openSMILE toolkit, selecting only the ComParE2016 features. We conducted a comparative study between various machine learning models like random forest (RF), support vector machine (SVM) and time-delayed neural network (TDNN) using both features. Our findings indicated that the stutter class outperformed the no stutter class by a significant margin for each classifier in binary classification. This prompted us to conduct two experiments: one with all classes and another with only stutter classes. The experimental analysis revealed an average F1-score with del-del MFCC features of 85.38% with RF, 90.10% with SVM when all classes are considered. Furthermore, 93.29% with RF, 93.71% with SVM when only stutter classes are considered. On the contrary, when we use ComParE2016 features we achieve an average accuracy of 86.79% with RF and 87.77% with SVM considering all classes. Whereas, 95.10% with RF, 94.34% with SVM when only stutter classes are considered. When TDNN was used with del-del MFCC, we achieved an accuracy of 79.02% and 92% with ComParE2016, implying better performance because of better feature extraction. It is noteworthy that throughout all the experiments, sound repetition is the best performing class indicating consistency in the learning process. The results achieved through this sampling technique improve the stutter detection and classification by a significant margin.