Dysarthria speakers face challenges in automatic speech recognition (ASR) due to reduced articulation, speech variability, atypical prosody, acoustic distortions, and limited training data. These difficulties impact accuracy and intelligibility, requiring personalized approaches for improved Dysarthric Speech Recognition (DSR). We propose a DSR approach using acoustic models and data augmentation techniques in life-long learning. Our approach combines Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM) for force alignment, and Time Delay Neural Network with Factorized layer (TDNN-F) in a multi-stream structure for DSR. We enhance system performance using self-attention and employ data augmentation techniques like pitch shifting, speed perturbation, noise injection, spec-augmentation, and generative adversarial network-based impaired speech synthesis. Experiments were conducted using a dysarthria dataset comprising individuals with Cerebral Palsy (CP), a prevalent cause of dysarthria. Our findings demonstrate the superior performance of our speech recognition system compared to ASR system trained on normal speech data. The system was evaluated in three scenarios: initial, personalized training with previous data, and personalized training with both previous data and synthesized dysarthric speech. Remarkably, personal training with previous data significantly enhanced system performance, particularly for individuals with severe CP. Additionally, combining personal training with previous data and synthesized dysarthric speech resulted in improved early-stage system performance, especially for individuals with severe CP. The severity of dysarthria, closely tied to the severity of CP, presents challenges for individuals with severe CP in the context of speech recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dysarthric Speech Recognition System with Life-Long Learning Technique in Taiwanese Mandarin

  • Alim Misbullah,
  • Hsiu-Wei Yeh,
  • Chia-Yuan Chang

摘要

Dysarthria speakers face challenges in automatic speech recognition (ASR) due to reduced articulation, speech variability, atypical prosody, acoustic distortions, and limited training data. These difficulties impact accuracy and intelligibility, requiring personalized approaches for improved Dysarthric Speech Recognition (DSR). We propose a DSR approach using acoustic models and data augmentation techniques in life-long learning. Our approach combines Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM) for force alignment, and Time Delay Neural Network with Factorized layer (TDNN-F) in a multi-stream structure for DSR. We enhance system performance using self-attention and employ data augmentation techniques like pitch shifting, speed perturbation, noise injection, spec-augmentation, and generative adversarial network-based impaired speech synthesis. Experiments were conducted using a dysarthria dataset comprising individuals with Cerebral Palsy (CP), a prevalent cause of dysarthria. Our findings demonstrate the superior performance of our speech recognition system compared to ASR system trained on normal speech data. The system was evaluated in three scenarios: initial, personalized training with previous data, and personalized training with both previous data and synthesized dysarthric speech. Remarkably, personal training with previous data significantly enhanced system performance, particularly for individuals with severe CP. Additionally, combining personal training with previous data and synthesized dysarthric speech resulted in improved early-stage system performance, especially for individuals with severe CP. The severity of dysarthria, closely tied to the severity of CP, presents challenges for individuals with severe CP in the context of speech recognition.