Using Embeddings of Pre-trained Models for Cross-Database Dysarthria Detection: Supervised vs. Self-supervised Approach
摘要
One of the main obstacles facing the research in the automatic detection of voice pathology is the lack of data. In this work, we try to tackle this problem by the use of self-supervised, and weakly supervised, pre-trained models, as feature extractors. We also perform cross-database experiments to test the ability of the proposed system to generalise across databases. UA-Speech and TORGO databases are used to carry the experiments, and the results are compared to baseline features. Two classifiers were used: SVM and a neural network. The pre-trained models achieved the best results on both databases when the neural network was used as classifier.