<p>Dysarthria is a neurological speech disorder characterized by articulatory impairment due to muscle weakness. Objective automated detection and severity classification of dysarthria enables timely intervention and tailored clinical management. Here, we propose a novel hybrid deep learning model that integrates Convolutional Neural Networks (CNN), Bidirectional Long Short-Term Memory (BiLSTM), and Transformer architectures and employs a unique cross-attention mechanism to fuse wavelet-based scalogram images with seven acoustic features, including Mel-Frequency Cepstral Coefficients and spectral descriptors. Evaluated on two public datasets, TORGO and UA Speech, the model achieves state-of-the-art accuracies of 98.74% and 99.86% for binary dysarthria detection, and 95.69% and 97.91% for multi-class severity classification, respectively. Cross‑dataset testing and tenfold cross‑validation confirm robustness; a paired t‑test across folds shows that the fused inputs significantly outperform an MFCC baseline (p &lt; 0.01). The results demonstrate that compared with previous methods,&#xa0;multimodal cross-attention feature fusion significantly enhances the detection of subtle dysarthric cues . This approach advances the accuracy and clinical applicability of automated speech disorder assessment, supporting early diagnosis and personalized treatment planning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Cross-Attentive CNN-BiLSTM-Transformer Network for Dysarthria Severity Classification

  • M. S. Remya,
  • Prakash Ishwar,
  • Prema Nedungadi

摘要

Dysarthria is a neurological speech disorder characterized by articulatory impairment due to muscle weakness. Objective automated detection and severity classification of dysarthria enables timely intervention and tailored clinical management. Here, we propose a novel hybrid deep learning model that integrates Convolutional Neural Networks (CNN), Bidirectional Long Short-Term Memory (BiLSTM), and Transformer architectures and employs a unique cross-attention mechanism to fuse wavelet-based scalogram images with seven acoustic features, including Mel-Frequency Cepstral Coefficients and spectral descriptors. Evaluated on two public datasets, TORGO and UA Speech, the model achieves state-of-the-art accuracies of 98.74% and 99.86% for binary dysarthria detection, and 95.69% and 97.91% for multi-class severity classification, respectively. Cross‑dataset testing and tenfold cross‑validation confirm robustness; a paired t‑test across folds shows that the fused inputs significantly outperform an MFCC baseline (p < 0.01). The results demonstrate that compared with previous methods, multimodal cross-attention feature fusion significantly enhances the detection of subtle dysarthric cues . This approach advances the accuracy and clinical applicability of automated speech disorder assessment, supporting early diagnosis and personalized treatment planning.