An Ensemble-Based Extra Feature Selection Approach for Predicting Heart Disease
摘要
Heart disease is a severe condition that has a significant impact on human life and is the leading cause of death in many countries. Clinicians can access clinical datasets to assist in diagnosing cardiovascular conditions and reduce the spread of this disease. Understanding these datasets and their patterns is essential in predicting the illness accurately. However, the proliferation of medical datasets in the era of big data has posed challenges for medical professionals in identifying the most relevant attributes for predicting heart disease. Therefore, this study aims to identify the most distinguishing characteristics within a high-dimensional dataset that will facilitate the accurate classification of patients with heart disease and non-patients with minimal complexity. The study conducted an experimental assessment of model performance generated through different classification algorithms, using both relevant features selected and the complete features set. An ensemble-based extra tree feature selection approach was employed for this purpose. The experiment utilized four datasets from reputable sources, employing various machine learning classification models with both complete and reduced feature subsets as inputs for the analysis. The evaluation of the classification models was based on accuracy, precision, and recall. The proposed feature selection technique achieved the highest classification accuracy, precision, and recall, reaching 97% when utilizing the extra tree classifier algorithm. These results underscore the potential of the feature selection method in identifying the most significant, related, and well-represented features for the target class, thereby enhancing classification performance.