Machine Learning Models Based Stuttering Classification
摘要
Speech disfluency refers to any interruption in the flow of spoken language caused by the speaker. There is a lack of data with regards to stuttering that is publicly available that can be used. We have used two datasets—UClass and LibriStutter which we used to apply different machine learning algorithms. We have used four different machine learning algorithms—K-Nearest Neighbors, Random Forest, Decision Trees, and Naive Bayes. Due to limited amount of data, machine learning techniques are utilized instead of deep learning techniques. There is a need to classify stuttering into different types as it helps to identify the different stutter types, develop personalized therapies and advance research initiatives for the classification of stuttering. Our classifier achieved an accuracy of 88.87% on the UClass dataset and 88.43% on the LibriStutter dataset. When equal samples were taken from both the UClass and LibriStutter datasets, an accuracy of 64.66% was achieved.