错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning for Predicting Stroke Occurrences Using Imbalanced Data

  • Nataliia Melnykova,
  • Yurii Patereha,
  • Liubomyr-Oleksii Chereshchuk,
  • Dariusz Sala

摘要

The research paper focused on predicting the occurrence of strokes, a severe threat to individuals’ health and lives. To address this, machine learning models were constructed using a highly imbalanced dataset, posing a significant challenge to the research. The processed data was used to build and compare different machine learning models for the stroke prediction classification problem by employing various data preprocessing techniques. Among the models tested, the random forest model emerged as the most successful, achieving precision, recall, and f1-score levels of 90%. Additionally, an accuracy metric of 90% was utilized. However, a random forest classifier was trained with optimal hyperparameters obtained through grid search to highlight the accuracy limitations in classification tasks. This model was trained on balanced data, illustrating the impracticality of relying solely on accuracy. As a result, the model exhibited a high accuracy of 96%. The findings of this study hold significance as they can aid healthcare professionals in implementing more effective preventive measures, thereby enhancing the likelihood of saving patients’ lives and preserving their health against strokes.