<p>Heart disease remains a leading cause of mortality worldwide, necessitating advancements in early prediction and diagnosis. While existing studies have employed machine learning (ML) techniques, they often lack specificity in addressing critical gaps, such as interpretability, computational efficiency, and performance on diverse datasets. This study aims to bridge these gaps by systematically evaluating traditional and advanced ML models, including Logistic Regression (LR), Support Vector Machine (SVM), Quadratic Discriminant Analysis (QDA), Naive Bayes (NB), and Fuzzy C-Means. The model was developed utilizing the Heart Disease Dataset by John Smith from Kaggle. Data preprocessing involved normalization and standardization, followed by feature selection based on XGBoost algorithm to enhance the model’s relevance and reduce dimensionality. The dataset was split into 75% for training and 25% for testing. Model performance was evaluated using metrics: Accuracy, Precision, Recall, Specificity, and F1 Score. Our proposed Fuzzy C-Means hybrid approach achieved performance improvement over baseline models. Specifically, the integration of Fuzzy C-Mean with Naïve Bayes improved accuracy by 8.34%, with K-Nearest Neighbors (K-NN) and Random Forest by 8.34% and 12.38%, respectively, and with Decision Tree by a notable 18.96%, achieving a highest accuracy of 99.22%. Our findings highlight improvements in prediction accuracy, interpretability, and scalability, demonstrating the practical significance of clustering-augmented ML techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Heart Disease Forecasting: Bridging Gaps in Interpretability, Efficiency, and Scalability Using Machine Learning

  • Muhammad Zohaib Khan,
  • Sarmad Ahmed Shaikh,
  • Abdullah Ayub Khan,
  • Aisha Imroz,
  • Muqaddas Salahuddin,
  • Priha Bhatti,
  • Shehzeen Dua Bhatti

摘要

Heart disease remains a leading cause of mortality worldwide, necessitating advancements in early prediction and diagnosis. While existing studies have employed machine learning (ML) techniques, they often lack specificity in addressing critical gaps, such as interpretability, computational efficiency, and performance on diverse datasets. This study aims to bridge these gaps by systematically evaluating traditional and advanced ML models, including Logistic Regression (LR), Support Vector Machine (SVM), Quadratic Discriminant Analysis (QDA), Naive Bayes (NB), and Fuzzy C-Means. The model was developed utilizing the Heart Disease Dataset by John Smith from Kaggle. Data preprocessing involved normalization and standardization, followed by feature selection based on XGBoost algorithm to enhance the model’s relevance and reduce dimensionality. The dataset was split into 75% for training and 25% for testing. Model performance was evaluated using metrics: Accuracy, Precision, Recall, Specificity, and F1 Score. Our proposed Fuzzy C-Means hybrid approach achieved performance improvement over baseline models. Specifically, the integration of Fuzzy C-Mean with Naïve Bayes improved accuracy by 8.34%, with K-Nearest Neighbors (K-NN) and Random Forest by 8.34% and 12.38%, respectively, and with Decision Tree by a notable 18.96%, achieving a highest accuracy of 99.22%. Our findings highlight improvements in prediction accuracy, interpretability, and scalability, demonstrating the practical significance of clustering-augmented ML techniques.