Pathological Speech and Electroglottography Signals Analysis Using Invariance Scattering Network
摘要
In the past few years, researchers have shown their interest in detecting speech disorders by analyzing pathological speeches; however, getting a robust approach to analyzing pathological speech is a complex and challenging problem. This paper aims to introduce a new multi-modal architecture that integrates speech and electroglottography (EGG) signals, exploring their potential in the automatic classification of healthy and pathological speech. The proposed framework consists of two parallel Invariance Scattering Networks (ISN), with one for speech signals and the other for EGG signals to extract scattering coefficients. These coefficients are then combined to form a more comprehensive feature set. Finally, a Support Vector Machine (SVM) classifier is employed to classify the healthy and pathological speech. Further, to evaluate the performance of the proposed system, several experiments were carried out using datasets from the Saarbruecken Voice Database (SVD). The results indicate that the proposed voice pathology detection approach reaches an accuracy of up to 85% when utilizing both speech and EGG samples. This paper also presents a comparative study of proposed method with other popular methods such as convolutional neural network (CNN) and openSMILE feature, and the findings show that the proposed approach outperforms state-of-the-art techniques in terms of accuracy and robustness.