Addressing Imbalanced Data in Stroke Prediction: An Oversampling Approach for Improved Accuracy
摘要
Stroke, a critical medical condition, demands accurate prediction for early intervention and improved patient outcomes. However, imbalanced datasets pose challenges in stroke prediction, with the minority class (stroke occurrences) often outnumbered by non-stroke cases. Our study introduces a novel approach to address this imbalance using oversampling techniques. We demonstrate that oversampling enhances machine learning model performance, improving accuracy. Analyzing diverse medical and demographic features, our approach successfully rebalances the dataset, allowing more accurate identification of individuals at stroke risk. We evaluate our approach on a real-world dataset, revealing significant accuracy improved. This research emphasizes the importance of addressing imbalanced data in stroke prediction and presents a promising oversampling technique for enhancing predictive models, ultimately improving patient care. It also highlights the potential for similar approaches in addressing imbalanced data challenges across healthcare domains.