A Novel Hybrid Deep Learning Model Optimized by Pelican Algorithm for Robust Dysarthric Speech Classification with Multi-feature Fusion and Attention-GAN
摘要
Dysarthria is a significant voice disorder that significantly hinders automatic human–machine interaction systems due to the low intelligibility of the voice. Over the last decade, various deep learning (DL)-based dysarthric speech classification (DSC) methods have been proposed, demonstrating superior performance compared to traditional machine learning (ML)-based DSC methods. However, the effectiveness of the DSC system is limited because of class imbalance problems, lower temporal depiction, inferior long-term dependency, and complex DL frameworks. This work presents DSC based on the novel improved pelican optimized BDGNet (IPOBDGNet), which combines Bidirectional Long Short-Term Memory (BiLSTM), Deep Convolutional Neural Network (DCNN), and Gated Recurrent Unit (GRU) and Multiple Dysarthric Features (MDFs). The improved pelican optimization algorithm (IPOA) is utilized for hyperparameter tuning such as learning rate, dropout, momentum and batch size of the BDGNet to enhance training performance and provide stability in training. Furthermore, it utilizes the attention-based Generative Adversarial Network (GAN) for augmenting speech data. It minimizes the class imbalance problem while retaining the synthetic voice’s perceptual quality, vocal characteristics, and prosodic features. The proposed IPOBDGNet, combined with an attention-based GAN, achieves an improved accuracy of 98.35% and 98.67% for the UASpeech and TORGO datasets, respectively, surpassing the traditional state-of-the-art for DSC methods.