A Hybrid Cross-Attentive CNN-BiLSTM-Transformer Network for Dysarthria Severity Classification
摘要
Dysarthria is a neurological speech disorder characterized by articulatory impairment due to muscle weakness. Objective automated detection and severity classification of dysarthria enables timely intervention and tailored clinical management. Here, we propose a novel hybrid deep learning model that integrates Convolutional Neural Networks (CNN), Bidirectional Long Short-Term Memory (BiLSTM), and Transformer architectures and employs a unique cross-attention mechanism to fuse wavelet-based scalogram images with seven acoustic features, including Mel-Frequency Cepstral Coefficients and spectral descriptors. Evaluated on two public datasets, TORGO and UA Speech, the model achieves state-of-the-art accuracies of 98.74% and 99.86% for binary dysarthria detection, and 95.69% and 97.91% for multi-class severity classification, respectively. Cross‑dataset testing and tenfold cross‑validation confirm robustness; a paired t‑test across folds shows that the fused inputs significantly outperform an MFCC baseline (p < 0.01). The results demonstrate that compared with previous methods, multimodal cross-attention feature fusion significantly enhances the detection of subtle dysarthric cues . This approach advances the accuracy and clinical applicability of automated speech disorder assessment, supporting early diagnosis and personalized treatment planning.