Unveiling Vocal Biomarkers: Investigating Parkinson’s Disease Detection Through PCA and Optimized MLP Models on Voice Datasets
摘要
Parkinson's disease (PD) is a neurodegenerative disorder related to dopamine enzyme causing tremors, rigidity, and cognitive impairment. This research study focuses on understanding its mechanisms, developing diagnostic tools, and exploring therapeutic interventions to improve the quality of life for affected individuals. Voice analysis in PD research aims to develop non-invasive diagnostic methods by examining vocal characteristics like pitch and rhythm, potentially serving as a biomarker for disease identification and monitoring. Advanced machine learning algorithms are being used in PD research to analyze diverse datasets and improve diagnostic accuracy. This study has collected 1348 voice data from 68 PD patients and 46 non-PD participants, with 1200 un-damaged recordings that are standardized and annotated with participant details. The dataset is split into 85% for training and 15 for testing. This research study explores the identification of PD using a novel combination of Principal Component Analysis and optimized Multilayer Perceptron models on the collected voice dataset. The proposed methodology for PD identification, based on PCA (15 dimensions) + PSO (5 – the size of swarm) + MLP (20 20), achieved an AUC of 0.97 and an accuracy of 0.9123, demonstrating its effectiveness. The PCA + MLP model performs AUC-0.92. The models PCA + PSO + MLP (10 10) and MLP (15 15) with AUCs of 0.93 and 0.96, respectively, and CA values of 0.8772 and 0.8947, which can be improved by increasing the MLP layer size.