<p>Cardiovascular disease or CVD is one of the leading causes of death in the modern world. The ability to forecast CVD enables healthcare practitioners to make wise decisions about their patient’s health. In the medical field, prediction models based on machine learning offer a superior method for assisting with patient health diagnoses. The most commonly used benchmark datasets in heart disease prediction are the Hungarian, Switzerland, Cleveland, and Long Beach datasets, each containing 303 instances with missing values in their features. The presence of missing values in the datasets can reduce the accuracy of the prediction model in predicting CVD accurately at an earlier stage. To overcome this limitation, in this study, a hybrid dataset is created by combining these datasets, resulting in a dataset with 920 instances and 14 attributes, all of which have null values. A correlation-based <i>GrpMean</i> algorithm is proposed to impute the missing values. This process creates a new dataset known as the <i>Hywin</i> (Hybrid Without Null values) dataset, which finally comprises 920 instances and 14 attributes with no null values. The proposed <i>GrpMean</i> imputation algorithm is then compared with existing imputation techniques such as KNNI, mean, mode, and median imputation. As part of the data pre-processing, feature scaling techniques are applied to normalize the attribute values. Subsequently, four machine learning classifiers, namely K-Nearest Neighbor (KNN), Random Forest (RF), Support Vector Machine (SVM), and CatBoost, are used for the earlier prediction of CVD. The experimental results demonstrated that the Random Forest classifier outperforms the other three classifiers, particularly when using the proposed <i>GrpMean</i> imputation. In comparison, the KNN classifier achieved 93.37%, the CatBoost classifier achieved 93.59% and the SVM classifier achieved 94.24% accuracy. Random Forest is further validated through stratified <i>sevenfold</i> cross-validation, and its accuracy ranged from 91.60 to 99.23%, with a mean accuracy of 94.68% which is best among the existing state-of-the-art techniques in the earlier prediction of the CVD.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GrpMean: predicting cardiovascular disease at an early stage

  • Hutashan Vishal Bhagat,
  • Saurabh Sharma,
  • Praveen Uppari,
  • Malapati Sruthi Laya Reddy

摘要

Cardiovascular disease or CVD is one of the leading causes of death in the modern world. The ability to forecast CVD enables healthcare practitioners to make wise decisions about their patient’s health. In the medical field, prediction models based on machine learning offer a superior method for assisting with patient health diagnoses. The most commonly used benchmark datasets in heart disease prediction are the Hungarian, Switzerland, Cleveland, and Long Beach datasets, each containing 303 instances with missing values in their features. The presence of missing values in the datasets can reduce the accuracy of the prediction model in predicting CVD accurately at an earlier stage. To overcome this limitation, in this study, a hybrid dataset is created by combining these datasets, resulting in a dataset with 920 instances and 14 attributes, all of which have null values. A correlation-based GrpMean algorithm is proposed to impute the missing values. This process creates a new dataset known as the Hywin (Hybrid Without Null values) dataset, which finally comprises 920 instances and 14 attributes with no null values. The proposed GrpMean imputation algorithm is then compared with existing imputation techniques such as KNNI, mean, mode, and median imputation. As part of the data pre-processing, feature scaling techniques are applied to normalize the attribute values. Subsequently, four machine learning classifiers, namely K-Nearest Neighbor (KNN), Random Forest (RF), Support Vector Machine (SVM), and CatBoost, are used for the earlier prediction of CVD. The experimental results demonstrated that the Random Forest classifier outperforms the other three classifiers, particularly when using the proposed GrpMean imputation. In comparison, the KNN classifier achieved 93.37%, the CatBoost classifier achieved 93.59% and the SVM classifier achieved 94.24% accuracy. Random Forest is further validated through stratified sevenfold cross-validation, and its accuracy ranged from 91.60 to 99.23%, with a mean accuracy of 94.68% which is best among the existing state-of-the-art techniques in the earlier prediction of the CVD.