Speaker independent dysarthria severity classification using synthesis-based augmentation
摘要
Dysarthria is a speech disorder caused by damage to the motor areas of the brain, resulting in speech that is effortful, slow, slurred, or disordered in rhythm and intonation. Assessing the severity of dysarthria helps pathologists plan therapy, track patient progress, and support automatic speech recognition systems. We approached the task of classifying dysarthria severity using deep learning models, including deep neural networks (DNN), convolutional neural networks (CNN), and long short-term memory networks (LSTM), applied to the standard UA Speech and TORGO databases with Mel-frequency cepstral coefficients (MFCCs). We also propose a data augmentation strategy based on multi-voice synthesis from lyrics. In the front end, a text-to-speech (TTS) converter generates synthetic speech in the target speaker’s voice. This is then processed by a style-transfer module to adjust the style of the synthesized speech. The best-performing DNN-MFCC framework achieved 97.68% and 97.16% accuracy in speaker-dependent scenarios for the UA Speech and TORGO databases, respectively. In speaker-independent scenarios, the accuracies were 40.12% and 48.07%, respectively.