Self-supervised Learning for Pathological Speech Analysis of Parkinson’s Disease
摘要
We explore self-supervised learning (SSL) techniques to predict Parkinson’s Disease (PD) by analysing PD speech signals. PD speech abnormalities can be identified up to five years before a medical diagnosis; however, a large labelled dataset is required to train a machine learning (ML) model to predict the disease. SSL can reduce data labelling requirements by learning data representations without explicit supervision. We focus on examining different SSL techniques. Tabular data with multiple acoustic and mel-spectrogram images are extracted from PD speech data to study the PD classification problem. Various SSL techniques, including TabNet, Encoder, Autoencoder, and SimCLR, are trained to classify the disease and compared with supervised ML methods. We observed that on several experimental scenarios, SSL performed better than supervised ML on the PD speech data, and even when only a third of the labelled data was SSL, accuracy only dropped by 1.86% on average. SSL approaches on tabular data performed much better than on mel-spectrogram image classification tasks, with the highest accuracy recorded at 94.15%.