Cardiovascular diseases are a leading cause of mortality, leading to an estimated 18 million fatalities globally. Elevated levels of biomarkers like cholesterol and glucose are critical in determining the associated risk. Prior studies have demonstrated the ability of machine learning in predicting these biomarkers, albeit with individual methods and with limited features. Hence, gaps remain in the exploration of ensemble methods and their efficiency in biomarker prediction. The study employs a range of machine learning models to predict these biomarkers. The models are trained and tested on three different data splits to evaluate their performance under varying data availability. A data driven approach using ensemble techniques is also proposed to enhance the prediction. We focused on three prediction scenarios: cholesterol levels alone, glucose levels alone, and the combined prediction of both cholesterol and glucose levels. Results show that Gradient Boosting Machine and Random Forest consistently provide highly accurate predictions across the three prediction scenarios. However, ensemble methods employed outperform the individual classifiers, achieving the highest performance across all scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Biomarker Prediction for Cardiovascular Health Using Data-Driven Methods

  • Anoushka Duggal,
  • Shivanjali Shaily,
  • Jyoti Maggu

摘要

Cardiovascular diseases are a leading cause of mortality, leading to an estimated 18 million fatalities globally. Elevated levels of biomarkers like cholesterol and glucose are critical in determining the associated risk. Prior studies have demonstrated the ability of machine learning in predicting these biomarkers, albeit with individual methods and with limited features. Hence, gaps remain in the exploration of ensemble methods and their efficiency in biomarker prediction. The study employs a range of machine learning models to predict these biomarkers. The models are trained and tested on three different data splits to evaluate their performance under varying data availability. A data driven approach using ensemble techniques is also proposed to enhance the prediction. We focused on three prediction scenarios: cholesterol levels alone, glucose levels alone, and the combined prediction of both cholesterol and glucose levels. Results show that Gradient Boosting Machine and Random Forest consistently provide highly accurate predictions across the three prediction scenarios. However, ensemble methods employed outperform the individual classifiers, achieving the highest performance across all scenarios.