Automatic Classification of Parkinson’s Disease Using Wav2vec Embeddings at Phoneme, Syllable, and Word Levels
摘要
Parkinson’s disease (PD) is a neurological condition that produces several speech deficits, typically known as hypokinetic dysarthria. PD involves motor impairments and muscle dysfunction in the phonatory apparatus, producing anomalies in oral communication. Speech signals have been used as a biomarker for diagnosis and monitoring of PD. In this work, we discriminate between PD patients and healthy controls based on patterns extracted from speech signals collected from Colombian spanish speakers, considering three different granularity levels: phoneme, syllable, and word. The Wav2vec 2.0 model is used to obtain frame-level representations of each utterance. These representations are grouped according to each granularity level using different statistical functionals. Each granularity level was evaluated independently, obtaining accuracies of \(86\%\) , \(80\%\) , and \(83\%\) for phonemes, syllables, and words, respectively. In addition, we identified the phonological classes with better discrimination capability. Nasals, approximant, and plosive classes were the three most accurate. We believe that this work constitutes a step forward in the development of automatic systems that support speech and language therapy of PD patients. For future work, we plan to model co-articulation information in words and syllables.