<p>Obesity is a worldwide epidemic posing a significant threat to the general population in terms of chronic metabolic diseases and low longevity. Clinical intervention and preventive health care require the early identification and proper risk prioritization. This paper compares six machine learning models, namely, Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, and Gradient Boosting using a publicly available obesity dataset in Mexico, Peru, and Colombia (n = 2111, 17 attributes). To overcome this issue of the dominance of certain classes, the Synthetic Minority Oversampling Technique (SMOTE) was used during the pre-processing to give fair consideration of the model. Gradient Boosting was the best with highest 95.93 percent accuracy, 0.96 precision, 0.96 recall and 0.96 F1-score among all the models. SHAP (SHapley Additive Explanations) analysis was used in increasing the explain ability with each predictor indicating the most significant determinant; frequent high-calorie food consumption (FAVC, |human|) (0.18), physical activity frequency (FAF, 0.15), family history of overweight (0.12), and hydration level (0.09) significance. These quantitative insights do not only enhance interpretability, but also conform to clinical knowledge of behavioral and genetic risk factors. The suggested interpretable framework can provide a strong and clear background to the design of data-driven decision-support tools that can be used to prevent obesity, develop individual counselling, and design the policies needed by the population.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Machine Learning Models for Obesity Classification with SHAP-Based Interpretability

  • Vaibhav Gandhi,
  • Yogesh Chaudhari,
  • Ajay Kumar,
  • Hitarth Revakar,
  • Ankit D. Oza,
  • Saneh Lata Yadav

摘要

Obesity is a worldwide epidemic posing a significant threat to the general population in terms of chronic metabolic diseases and low longevity. Clinical intervention and preventive health care require the early identification and proper risk prioritization. This paper compares six machine learning models, namely, Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, and Gradient Boosting using a publicly available obesity dataset in Mexico, Peru, and Colombia (n = 2111, 17 attributes). To overcome this issue of the dominance of certain classes, the Synthetic Minority Oversampling Technique (SMOTE) was used during the pre-processing to give fair consideration of the model. Gradient Boosting was the best with highest 95.93 percent accuracy, 0.96 precision, 0.96 recall and 0.96 F1-score among all the models. SHAP (SHapley Additive Explanations) analysis was used in increasing the explain ability with each predictor indicating the most significant determinant; frequent high-calorie food consumption (FAVC, |human|) (0.18), physical activity frequency (FAF, 0.15), family history of overweight (0.12), and hydration level (0.09) significance. These quantitative insights do not only enhance interpretability, but also conform to clinical knowledge of behavioral and genetic risk factors. The suggested interpretable framework can provide a strong and clear background to the design of data-driven decision-support tools that can be used to prevent obesity, develop individual counselling, and design the policies needed by the population.