Background <p>This investigation aimed to evaluate the relationship between a combined triglyceride-glucose (TyG)–adiposity index and carotid plaque within a low-income rural cohort, and to apply machine-learning models alongside SHapley Additive exPlanations (SHAP) for detailed interpretation.</p> Methods <p>We conducted a cross-sectional analysis of 1,960 adults enrolled from documented low-income rural areas. Sociodemographic variables, lifestyle habits, anthropometric indices, and biochemical markers were systematically recorded. A binary logistic model served to predict carotid plaque, while restricted cubic splines (RCS) examined possible non-linear associations between the TyG–ABSI composite (triglyceride-glucose index combined with a body-shape index) and plaque risk. Next, ten machine-learning classifiers—Logistic Regression, Support Vector Machine (SVM), Gradient Boosting Machine (GBM), Neural Network, Random Forest, XGBoost, K-Nearest Neighbors (KNN), AdaBoost, LightGBM, and CatBoost—were trained and internally validated. Hierarchical 5-fold cross-validation was implemented in the training set to fine-tune the hyperparameters. Model performance was evaluated on both the training and validation sets using accuracy, sensitivity, specificity, precision, F1 score, and the area under the ROC curve (AUC). SHAP values were computed to quantify and visualize feature contributions.</p> Results <p>Among 1,960 participants, the overall prevalence of carotid plaque was 48.3%, with 59.0% in men and 41.7% in women. In multivariable-adjusted logistic regression, sex, age, systolic blood pressure (SBP), and TyG-ABSI were independently associated with carotid plaque. Each one-unit increment in TyG-ABSI corresponded to a 20% higher odds of plaque presence (OR = 1.20; 95% CI 1.02–1.42; <i>P</i> = 0.025). RCS modelling revealed a non-linear relationship (P for non-linearity = 0.038), risk rose with TyG-ABSI up to = 7.75 and then declined, yielding an inverted-U trend around this inflection point. Among ML algorithms, logistic regression achieved the best generalization on the validation set (accuracy = 0.665, F1 = 0.58, AUC = 0.67). SHAP analysis confirmed the predictive importance of TyG-ABSI.</p> Conclusions <p>In this cross-sectional study of a low-income rural population, the TyG-ABSI index demonstrated a significant nonlinear relationship with carotid plaque risk. Among the machine learning algorithms evaluated, logistic regression achieved the highest predictive accuracy. SHAP-based visualizations further revealed the key features driving this distinction. Future research is needed to validate causal associations through prospective studies.</p> Clinical trial number <p>Not applicable.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TyG-ABSI as a novel metabolic obesity indicator for carotid plaque: an explainable machine learning study using SHAP in low-income population

  • Juan Hao,
  • Ran Chen,
  • Diliyaer Abudukeremu,
  • Xiao Li,
  • Yiwei Zhang,
  • Lifeng Wang,
  • Chenxi Fan,
  • Chunsheng Yang,
  • Xianjia Ning,
  • Jinghua Wang,
  • Yan Li

摘要

Background

This investigation aimed to evaluate the relationship between a combined triglyceride-glucose (TyG)–adiposity index and carotid plaque within a low-income rural cohort, and to apply machine-learning models alongside SHapley Additive exPlanations (SHAP) for detailed interpretation.

Methods

We conducted a cross-sectional analysis of 1,960 adults enrolled from documented low-income rural areas. Sociodemographic variables, lifestyle habits, anthropometric indices, and biochemical markers were systematically recorded. A binary logistic model served to predict carotid plaque, while restricted cubic splines (RCS) examined possible non-linear associations between the TyG–ABSI composite (triglyceride-glucose index combined with a body-shape index) and plaque risk. Next, ten machine-learning classifiers—Logistic Regression, Support Vector Machine (SVM), Gradient Boosting Machine (GBM), Neural Network, Random Forest, XGBoost, K-Nearest Neighbors (KNN), AdaBoost, LightGBM, and CatBoost—were trained and internally validated. Hierarchical 5-fold cross-validation was implemented in the training set to fine-tune the hyperparameters. Model performance was evaluated on both the training and validation sets using accuracy, sensitivity, specificity, precision, F1 score, and the area under the ROC curve (AUC). SHAP values were computed to quantify and visualize feature contributions.

Results

Among 1,960 participants, the overall prevalence of carotid plaque was 48.3%, with 59.0% in men and 41.7% in women. In multivariable-adjusted logistic regression, sex, age, systolic blood pressure (SBP), and TyG-ABSI were independently associated with carotid plaque. Each one-unit increment in TyG-ABSI corresponded to a 20% higher odds of plaque presence (OR = 1.20; 95% CI 1.02–1.42; P = 0.025). RCS modelling revealed a non-linear relationship (P for non-linearity = 0.038), risk rose with TyG-ABSI up to = 7.75 and then declined, yielding an inverted-U trend around this inflection point. Among ML algorithms, logistic regression achieved the best generalization on the validation set (accuracy = 0.665, F1 = 0.58, AUC = 0.67). SHAP analysis confirmed the predictive importance of TyG-ABSI.

Conclusions

In this cross-sectional study of a low-income rural population, the TyG-ABSI index demonstrated a significant nonlinear relationship with carotid plaque risk. Among the machine learning algorithms evaluated, logistic regression achieved the highest predictive accuracy. SHAP-based visualizations further revealed the key features driving this distinction. Future research is needed to validate causal associations through prospective studies.

Clinical trial number

Not applicable.