In recent decades, cardiovascular diseases have emerged as a leading cause of mortality worldwide, underscoring the critical need for accurate diagnostic and predictive methodologies. Among the various cardiovascular ailments, coronary artery disease and chronic heart failure are predominant causes of heart attacks, which substantially contribute to the high death rates. Traditionally, angiography has been the cornerstone for diagnosing heart disease; however, the advent of big data and advanced computational techniques offers a transformative potential for enhancing diagnostic accuracy and patient management. This paper strives for the application of machine learning algorithms in the predictive analysis of heart disease, exploring five distinct models: Logistic regression (LR), support vector machines (SVM), K-nearest neighbor (KNN), decision tree (DT), and random forest (RF). Among these, the decision tree model is highlighted for its superior precision in predicting cardiac illnesses. The integration of these computational models with big data analytics enables clinicians to systematically evaluate large datasets, thereby improving the prediction of heart disease risks and optimizing therapeutic strategies. Furthermore, the recent research introduces a newly compiled dataset from 2020, encompassing approximately 320,000 records with 18 features, to refine the diagnostic processes through advanced feature engineering and machine learning techniques. The utilization of the UCI dataset along with innovative computational models—namely logistic regression and two additional undisclosed methods—has demonstrated a better prediction accuracy using logistic regression than other classifiers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Heart Disease Prediction Using Machine Learning with Feature Engineering

  • Adnan Abass,
  • Gaurav Bathla,
  • Vikas Wasson

摘要

In recent decades, cardiovascular diseases have emerged as a leading cause of mortality worldwide, underscoring the critical need for accurate diagnostic and predictive methodologies. Among the various cardiovascular ailments, coronary artery disease and chronic heart failure are predominant causes of heart attacks, which substantially contribute to the high death rates. Traditionally, angiography has been the cornerstone for diagnosing heart disease; however, the advent of big data and advanced computational techniques offers a transformative potential for enhancing diagnostic accuracy and patient management. This paper strives for the application of machine learning algorithms in the predictive analysis of heart disease, exploring five distinct models: Logistic regression (LR), support vector machines (SVM), K-nearest neighbor (KNN), decision tree (DT), and random forest (RF). Among these, the decision tree model is highlighted for its superior precision in predicting cardiac illnesses. The integration of these computational models with big data analytics enables clinicians to systematically evaluate large datasets, thereby improving the prediction of heart disease risks and optimizing therapeutic strategies. Furthermore, the recent research introduces a newly compiled dataset from 2020, encompassing approximately 320,000 records with 18 features, to refine the diagnostic processes through advanced feature engineering and machine learning techniques. The utilization of the UCI dataset along with innovative computational models—namely logistic regression and two additional undisclosed methods—has demonstrated a better prediction accuracy using logistic regression than other classifiers.