Enhancing Dysarthria Detection: Harnessing Ensemble Models and MFCC
摘要
A dysarthria involves difficulties with articulation, respiration, resonance, phonation, and prosody. Dysarthria management and treatment can be guided by an early and accurate diagnosis. Speech therapists traditionally use perceptual assessments to diagnose, which can be subjective. An objective analysis of speech features can be performed using automated methods based on machine learning. An extensive publicly available dataset, the TORGO database, is used to classify dysarthric speech. From recordings of Parkinson's dysarthric and control subjects, melt-frequency cepstral coefficients (MFCCs) are extracted. Normalization of the data is performed, and training and test sets are separated. In this paper, the performance of artificial neural networks (ANNs), ensemble convolutional neural networks (CNNs), and ensemble artificial neural networks (ANNs) for dysarthria detection is evaluated. As a result of the research, it has been able to develop diagnostic and assessment tools for dysarthria, improving the quality of life for people with this condition and providing them with some hope for early intervention. This work highlights the promising potential of automated speech analysis for dysarthria detection by employing state-of-the-art machine learning techniques.