Pathological Voice Detection Based Deep Learning Techniques
摘要
Pathological voice is a feared problem for people due to its significant impact on their quality of life. This work introduces a novel feature extraction approach for voice disorder detection based on spectrograms and deep learning algorithms. The objective is to determine whether the speech is healthy or paralyzed. Audio voice signals were converted into spectrograms and fed into various pre-trained models, including MobileNet, VGG19, VGG16, and NASNetLarge. The models were trained using voice signals of the vowels /u/, /i/, and /a/ at low, neutral, and high pitch levels from the Saarbrucken Voice Dataset (SVD). The models were evaluated three times, first using neutral, then neutral with high voices, and finally using neutral, high, and low pitch levels together. Performance was evaluated using F1-score, recall, accuracy, and precision. The MobileNet and VGG19 achieved the best performance, reaching 100% on all performance metrics. The findings demonstrate the efficacy of deep learning techniques in identifying paralysis voices. They also illustrate the impact of including different pitch levels of speech, which results in poor performance compared to using a single pitch level in training and testing tasks.