Cardiac events rank among the leading causes of global mortality and contribute significantly to public health concerns. With an emphasis on important risk factors which are later on discussed, this study uses machine learning and classification approaches to predict the likelihood of cardiac stroke. The dataset, obtained from Kaggle, consists of 5,110 records with 12 variables. Several preprocessing techniques were employed, including resampling strategies like SMOTE and random undersampling to handle class imbalance, frequency and label encoding for categorical features, and median imputation for missing data. The performance of four machine learning models—ensemble, K-Nearest Neighbors (K-NN), neural networks, and Support Vector Machine (SVM)—was evaluated in terms of accuracy, precision, recall, specificity, and F-score. Among the models, the ensemble method delivered superior results, achieving the highest accuracy, precision, and F-score, positioning it as the most effective approach for stroke prediction in clinical applications. K-NN stood out in recall but had a lower precision, indicating a trade-off. Neural networks demonstrated balanced performance, adept at capturing complex, non-linear patterns, while SVM produced reliable outcomes with a strong equilibrium between precision and recall. These findings emphasize the importance of selecting the right model for heart attack prediction tasks, with the ensemble approach proving the most beneficial in practical scenarios. This research contributes to ongoing efforts to enhance early stroke detection, offering healthcare professionals a valuable tool for implementing preventative strategies and reducing stroke occurrences.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Stroke Prediction Through Machine Learning: A Comparative Study of Key Algorithms

  • Ayush Dixit,
  • Sneha Sharma,
  • Jayaprakash Vemuri

摘要

Cardiac events rank among the leading causes of global mortality and contribute significantly to public health concerns. With an emphasis on important risk factors which are later on discussed, this study uses machine learning and classification approaches to predict the likelihood of cardiac stroke. The dataset, obtained from Kaggle, consists of 5,110 records with 12 variables. Several preprocessing techniques were employed, including resampling strategies like SMOTE and random undersampling to handle class imbalance, frequency and label encoding for categorical features, and median imputation for missing data. The performance of four machine learning models—ensemble, K-Nearest Neighbors (K-NN), neural networks, and Support Vector Machine (SVM)—was evaluated in terms of accuracy, precision, recall, specificity, and F-score. Among the models, the ensemble method delivered superior results, achieving the highest accuracy, precision, and F-score, positioning it as the most effective approach for stroke prediction in clinical applications. K-NN stood out in recall but had a lower precision, indicating a trade-off. Neural networks demonstrated balanced performance, adept at capturing complex, non-linear patterns, while SVM produced reliable outcomes with a strong equilibrium between precision and recall. These findings emphasize the importance of selecting the right model for heart attack prediction tasks, with the ensemble approach proving the most beneficial in practical scenarios. This research contributes to ongoing efforts to enhance early stroke detection, offering healthcare professionals a valuable tool for implementing preventative strategies and reducing stroke occurrences.