This paper addresses the problem of identifying risk factors associated with diabetes using advanced machine learning techniques. The method used is based on combining rigorous data preparation with an exhaustive evaluation of multiple algorithms, optimizing predictive accuracy and facilitating the interpretation of the results. The development of the study is organized in three phases: Data preparation: From a public dataset, loading, detailed analysis of variables, denoising and data transformation are carried out. These steps ensure that the information is of high quality and ready for exploratory and predictive analysis. Classifier testing: Different machine learning algorithms are evaluated, from classical approaches to advanced methods such as J48, KNN, Linear Regression, Multi-Layer Perceptron (MLP), AdaBoost, XGBoost, CatBoost, Gradient Boosting, LightGBM and Random Forest. During this phase, exploratory and predictive analysis is performed to measure the performance of the methods based on seven key metrics. Selection of the best method: The results obtained allow us to identify the best performing method. In this case, Gradient Boosting and Random Forest proved to be the most efficient, while Multilayer Perceptron (MLP) presented the lowest performance in both the training and testing phases. This integrated approach not only ensures an efficient extraction of knowledge from the data, but also provides a detailed comparison of the performance of the methods, allowing to identify the most suitable to address this type of problem.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing a Robust Method for Diabetes Prediction with Machine Learning and Ensemble Models

  • Patricio Sizalima,
  • Paul Espinoza,
  • Remigio Hurtado,
  • Rodolfo Bojorque

摘要

This paper addresses the problem of identifying risk factors associated with diabetes using advanced machine learning techniques. The method used is based on combining rigorous data preparation with an exhaustive evaluation of multiple algorithms, optimizing predictive accuracy and facilitating the interpretation of the results. The development of the study is organized in three phases: Data preparation: From a public dataset, loading, detailed analysis of variables, denoising and data transformation are carried out. These steps ensure that the information is of high quality and ready for exploratory and predictive analysis. Classifier testing: Different machine learning algorithms are evaluated, from classical approaches to advanced methods such as J48, KNN, Linear Regression, Multi-Layer Perceptron (MLP), AdaBoost, XGBoost, CatBoost, Gradient Boosting, LightGBM and Random Forest. During this phase, exploratory and predictive analysis is performed to measure the performance of the methods based on seven key metrics. Selection of the best method: The results obtained allow us to identify the best performing method. In this case, Gradient Boosting and Random Forest proved to be the most efficient, while Multilayer Perceptron (MLP) presented the lowest performance in both the training and testing phases. This integrated approach not only ensures an efficient extraction of knowledge from the data, but also provides a detailed comparison of the performance of the methods, allowing to identify the most suitable to address this type of problem.