Machine Learning-Based Models for the Preemptive Diagnosis of Sickle Cell Anemia Using Clinical Data
摘要
Sickle cell anemia (SCA) is a chronic genetic condition that results in sickle-shaped red blood cells due to aberrant hemoglobin production. This condition results in impaired oxygen transport, vessel occlusion, and various complications, severely impacting patients’ health and quality of life. In Saudi Arabia, SCA presents a significant public health challenge, particularly in the Eastern Province, where prevalence rates are notably high. Early diagnosis and intervention are crucial for effective management and improved outcomes. Leveraging machine learning (ML) algorithms alongside traditional diagnostic methods offers promising avenues for enhancing SCA diagnosis and prediction. By integrating clinical data from blood samples with ML techniques, this study seeks to develop accessible and interpretable predictive models for preemptive SCA diagnosis. Drawing on a dataset collected from a pediatric hospital-based study in Sudan, encompassing sociodemographic variables and clinical parameters, this research aims to overcome prior limitations in ML-based SCA diagnosis, particularly the scarcity of studies incorporating clinical data. Through the application of various ML algorithms, including KNN, XGBoost, and Random Forest, this study endeavors to facilitate targeted interventions and personalized treatment plans, ultimately mitigating the burden of SCA in affected populations. Results indicate significant improvements in diagnostic accuracy with Random Forest demonstrating the highest accuracy at 86.77% using only 5 features, alongside 89% precision, 87% recall, and 83% f1-score.