Identification of Disfluency Among Children Using Efficient Machine Learning Techniques
摘要
Disfluency, which refers to any deviation from the anticipated fluency of spoken language, is a considerable issue affecting a substantial number of people in India. An estimated 11–12 million individuals in India suffer from stuttering, a neurodevelopmental disorder characterized by speech production disfluencies. Stuttering can have detrimental effects on various aspects of life, including education, career opportunities, and social interactions, affecting approximately 1% of the adult population and 5% of children. The current datasets are only suitable for adults, which makes them inappropriate for children. This research paper presents the methodology for creating a Telugu dataset (TLD-ISC) that caters to children aged 7–13 years from various socioeconomic backgrounds. The data samples were obtained at a frequency of 44,100 MHz with an average duration of 20 s, and each subject was recorded three times for consistency. The data was manually annotated, and MFCC features were extracted. Support Vector Machine (SVM) and K-Nearest Neighbour (KNN) algorithms were trained using TLD-ISC dataset. SVM achieved state of art results with 97.49% accuracy compared to the KNN classifier.