错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Logistic Regression, Random Forest, XGBoost, LightGBM, and Voting Classifier for Early Stroke Detection

  • Abha Sharma,
  • Saboor Uddin Ahmed,
  • Rabia Musheer Aziz,
  • Mohd Asif Shah

摘要

Stroke is a worldwide leading cause of death and long-term disability, making early prediction and prevention extremely crucial. The current research recommends machine learning-based stroke risk prediction through Logistic Regression, Random Forest, LightGBM, XGBoost, and Voting Classifier ensemble. Models were trained with publicly available data consisting of demographic, behavioral, and medical history features like age, hypertension, heart disease, type of work, smoking, and BMI. Data preprocessing involved class imbalance handling using Synthetic Minority Over-sampling Technique (SMOTE), normalization, and feature selection. Performance was measured mainly through confusion matrices to examine true and false classifications extensively. Voting Classifier, which compiled the advantages of the individual models, performed better than the individual classifiers in terms of balanced accuracy and stability. Logistic Regression gave us a good baseline, Random Forest provided interpretability via feature importance, and gradient boosting models (LightGBM and XGBoost) provided improved precision. The ensemble method ensured minimum overfitting and better generalizability, depicting the potential of machine learning—particularly ensemble methods—for developing robust and scalable stroke prediction systems to aid early clinical diagnosis and preventive care.