Machine Learning Approaches for COVID-19 Classification and Potential Novel Therapeutic Targets
摘要
The recent pandemic of Coronavirus disease (COVID-19) has profoundly impacted both the economic and healthcare sectors globally. The virus remains highly relevant in 2024, since it is still mutating and new variants are emerging. A key aspect of COVID-19 research is classification, such as disease severity classification or classification into positive and negative cases. Insights from classification tasks of high accuracy can be leveraged for management of the disease, drug development, and therapeutic approaches to improve patients’ health outcomes. The objective of this study was to compare three state-of-the-art supervised Machine Learning (ML) classifiers, (Random Forest (RF), Extreme Gradient Boosting (XGBoost) and Support Vector Machine (SVM)), in order to investigate how accurately they can discriminate between positive and negative COVID-19 groups according to their proteomic profile using publicly available high-dimensional protein abundance data with a small sample size (N = 52). The results demonstrated that the XGBoost model (accuracy 92%) has superior performance compared to the SVM model (accuracy 69%) and the RF model (accuracy 85%). This study contributes valuable evidence to the literature corpora, by reinforcing the great potential of ML algorithms for classification of COVID-19 in high-dimensional and small sample size data. Additionally, it uncovered several promising biomarker candidates identified by RF and XGBoost as highly contributing proteins, with a previously unexplored link to COVID-19. The proteins identified herein provide a solid basis for laboratory investigations, presenting significant potential as candidate biomarkers not only for COVID-19, but also for ‘Disease X,’ which may exhibit analogous biological mechanisms, such as proteomic processes and inflammatory pathways.