Handling Imbalanced Datasets with Real-World Positive Samples in Dengue Prediction Using Machine and Deep Learning Models
摘要
Dengue fever has been a growing concern for public health around the world, and we have been motivated to find better ways to diagnose it accurately. We decided to collect data on patients who tested positive for dengue, making sure to note down their symptoms. Then, we paired that with some other datasets we found on Kaggle on COVID-19, flu, and seasonal fever to use as our negative cases. After cleaning it and using SMOTE to balance out the numbers, we tested models like Logistic Regression, SVM, Random Forest, LightGBM, CatBoost, K-NN, XGBoost, MLP, LSTM, and TabNet. We focused on metrics such as accuracy, precision, recall, F1-score, and AUC to see how these models performed. Models such as Logistic Regression, SVM, Random Forest, LightGBM, CatBoost, K-NN, and LSTM performed well, reaching an accuracy of 96.43%. They also had perfect precision at 1.0, with a recall of 85.71% and an F1-score of 92.31%. LSTM did overall better and had the best AUC at 0.982, indicating it was effective at picking up patterns in the data. However, models like XGBoost and MLP still delivered an F1-score of 82.76%. For TabNet, it reached an F1-score of 31.58%. Simpler models sometimes surpassed the complex deep learning models. We think the imbalanced data, with fewer true dengue cases, posed challenges for models like TabNet in accurately identifying positives, underscoring the difficulty of managing imbalanced datasets.