Biomarker Prediction for Cardiovascular Health Using Data-Driven Methods
摘要
Cardiovascular diseases are a leading cause of mortality, leading to an estimated 18 million fatalities globally. Elevated levels of biomarkers like cholesterol and glucose are critical in determining the associated risk. Prior studies have demonstrated the ability of machine learning in predicting these biomarkers, albeit with individual methods and with limited features. Hence, gaps remain in the exploration of ensemble methods and their efficiency in biomarker prediction. The study employs a range of machine learning models to predict these biomarkers. The models are trained and tested on three different data splits to evaluate their performance under varying data availability. A data driven approach using ensemble techniques is also proposed to enhance the prediction. We focused on three prediction scenarios: cholesterol levels alone, glucose levels alone, and the combined prediction of both cholesterol and glucose levels. Results show that Gradient Boosting Machine and Random Forest consistently provide highly accurate predictions across the three prediction scenarios. However, ensemble methods employed outperform the individual classifiers, achieving the highest performance across all scenarios.