Benchmarking Machine Learning Models for Obesity Classification with SHAP-Based Interpretability
摘要
Obesity is a worldwide epidemic posing a significant threat to the general population in terms of chronic metabolic diseases and low longevity. Clinical intervention and preventive health care require the early identification and proper risk prioritization. This paper compares six machine learning models, namely, Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, and Gradient Boosting using a publicly available obesity dataset in Mexico, Peru, and Colombia (n = 2111, 17 attributes). To overcome this issue of the dominance of certain classes, the Synthetic Minority Oversampling Technique (SMOTE) was used during the pre-processing to give fair consideration of the model. Gradient Boosting was the best with highest 95.93 percent accuracy, 0.96 precision, 0.96 recall and 0.96 F1-score among all the models. SHAP (SHapley Additive Explanations) analysis was used in increasing the explain ability with each predictor indicating the most significant determinant; frequent high-calorie food consumption (FAVC, |human|) (0.18), physical activity frequency (FAF, 0.15), family history of overweight (0.12), and hydration level (0.09) significance. These quantitative insights do not only enhance interpretability, but also conform to clinical knowledge of behavioral and genetic risk factors. The suggested interpretable framework can provide a strong and clear background to the design of data-driven decision-support tools that can be used to prevent obesity, develop individual counselling, and design the policies needed by the population.