Early diagnosis of dengue poses a considerable challenge in resource limited settings. Dengue fever may progress to life threatening diseases and clinicians need to identify high risk patients before their condition worsen. A publicly available Vietnamese dataset of 2301 Vietnamese children aged 5 to 15 years who developed dengue shock syndrome (DSS) between 2001 and 2009 was used to build a DSS predictive model using machine learning algorithms. Missing values of clinical and laboratory variables were imputed using the K-Nearest Neighbor (KNN) technique. Subsequently, feature selection was performed using Recursive Feature Elimination with Cross Validation (RFECV) and the Chi-squared test. Selected parameters included age, weight, days of illness, temperature, hematocrit level, platelet count and history of vomiting as predictors of DSS. Random Under-Sampling (RUS) and Synthetic Minority Over-sampling (SMOTE) were used to address class imbalance in the dataset. To develop predictive models, Logistic Regression (LR), Random Forest (RF), XGBoost (XGB) and Support Vector Machine (SVM) were trained. The results showed that SMOTE combined with XGBoost and SVM classifiers achieved the highest performance, with a 95.97% F1-score for both classifiers and an area under receiver operating characteristic (AUROC) of 98.48% for XGBoost and 98.89% for SVM classifier. This research shows potential for developing a predictive model by comparing several machine learning approaches with SMOTE, aimed at the early detection of DSS in resource limited settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Machine Learning Algorithms to Predict Dengue Shock Syndrome

  • Sachanee Madhukala,
  • Sulanie Perera

摘要

Early diagnosis of dengue poses a considerable challenge in resource limited settings. Dengue fever may progress to life threatening diseases and clinicians need to identify high risk patients before their condition worsen. A publicly available Vietnamese dataset of 2301 Vietnamese children aged 5 to 15 years who developed dengue shock syndrome (DSS) between 2001 and 2009 was used to build a DSS predictive model using machine learning algorithms. Missing values of clinical and laboratory variables were imputed using the K-Nearest Neighbor (KNN) technique. Subsequently, feature selection was performed using Recursive Feature Elimination with Cross Validation (RFECV) and the Chi-squared test. Selected parameters included age, weight, days of illness, temperature, hematocrit level, platelet count and history of vomiting as predictors of DSS. Random Under-Sampling (RUS) and Synthetic Minority Over-sampling (SMOTE) were used to address class imbalance in the dataset. To develop predictive models, Logistic Regression (LR), Random Forest (RF), XGBoost (XGB) and Support Vector Machine (SVM) were trained. The results showed that SMOTE combined with XGBoost and SVM classifiers achieved the highest performance, with a 95.97% F1-score for both classifiers and an area under receiver operating characteristic (AUROC) of 98.48% for XGBoost and 98.89% for SVM classifier. This research shows potential for developing a predictive model by comparing several machine learning approaches with SMOTE, aimed at the early detection of DSS in resource limited settings.