<p>Parkinson’s disease (PD) is a neurodegenerative disorder (ND) with vocal impairments that complicate its differentiation from other neurological disorders. This study introduces a novel database of voice recordings from 40 PD patients and 20 patients with other neurological disorders, collected in both clinical and natural environments. A comprehensive set of 218 acoustic features was extracted, and machine learning algorithms—including Random Forest, Support Vector Machine, k-Nearest Neighbors, AdaBoost and others—were evaluated under three scenarios: no dimensionality reduction, linear Principal Component Analysis (PCA), and Non-linear Principal Component Analysis (NPCA). Using Leave-One-Subject-Out (LOSO) cross-validation, NPCA significantly enhanced classification performance, with algorithms achieving up to 95% accuracy. The optimal configuration, NPCA_5, demonstrated that reducing dimensionality effectively captures essential non-linear relationships while minimizing noise. Performance metrics, including accuracy, sensitivity, specificity, F1 score, and Matthews Correlation Coefficient (MCC), consistently improved with fewer NPCA components. The study highlights the importance of NPCA in capturing non-linear acoustic features and optimizing model performance. The diverse recording conditions and participant demographics enhance the dataset’s ecological validity, making it valuable for developing practical diagnostic tools. These findings advocate for the strategic use of NPCA in machine learning pipelines to differentiate PD from other ND.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning-based voice analysis for differentiating Parkinson’s disease and neurological disorders

  • Achraf Benba,
  • Sara Sandabad,
  • Fatima Ezzahra Mouas,
  • Latifa Doudach

摘要

Parkinson’s disease (PD) is a neurodegenerative disorder (ND) with vocal impairments that complicate its differentiation from other neurological disorders. This study introduces a novel database of voice recordings from 40 PD patients and 20 patients with other neurological disorders, collected in both clinical and natural environments. A comprehensive set of 218 acoustic features was extracted, and machine learning algorithms—including Random Forest, Support Vector Machine, k-Nearest Neighbors, AdaBoost and others—were evaluated under three scenarios: no dimensionality reduction, linear Principal Component Analysis (PCA), and Non-linear Principal Component Analysis (NPCA). Using Leave-One-Subject-Out (LOSO) cross-validation, NPCA significantly enhanced classification performance, with algorithms achieving up to 95% accuracy. The optimal configuration, NPCA_5, demonstrated that reducing dimensionality effectively captures essential non-linear relationships while minimizing noise. Performance metrics, including accuracy, sensitivity, specificity, F1 score, and Matthews Correlation Coefficient (MCC), consistently improved with fewer NPCA components. The study highlights the importance of NPCA in capturing non-linear acoustic features and optimizing model performance. The diverse recording conditions and participant demographics enhance the dataset’s ecological validity, making it valuable for developing practical diagnostic tools. These findings advocate for the strategic use of NPCA in machine learning pipelines to differentiate PD from other ND.