<p>In this research, we present an interpretable AutoML approach for the early diagnosis of hypertension and hyperinsulinemia among adolescents, conditions that are critical to identify during these formative years due to their requirement for lifelong care and monitoring. The dataset, collected from 2019 to 2022 by Serbia’s Healthcare Center through an observational cross-sectional study, posed challenges common to medical datasets, including imbalances, data scarcity, and a need for transparent, explainable predictive models. To counter these issues, we utilized three AutoML frameworks - AutoGluon, H2O, and MLJAR - in conjunction with a Tabular Variational Autoencoder (TVAE) to synthetically augment the data points, Prinicipal Component Analysis (PCA) for dimensionality reduction, and SHapley Additive exPlanations (SHAP) and Permutation feature importance analyses to extract insights from the results. AutoGluon outperformed the others on the original dataset, delivering better results with weighted ensemble models for both conditions under a 12-minute budget-time constraint and maintaining all evaluation metrics below a 4% threshold, all without the need for further scaling or calibration in the experimental setup. Our research underscores the broad applicability of the current AutoML paradigm, highlighting its particular benefits for the healthcare domain and diagnostics, where such advanced tools can enhance patient care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Innovations in early detection of chronic non-communicable diseases among adolescents through an easy-to-Use AutoML paradigm

  • Nevena Rankovic,
  • Dragica Rankovic,
  • Igor Lukic

摘要

In this research, we present an interpretable AutoML approach for the early diagnosis of hypertension and hyperinsulinemia among adolescents, conditions that are critical to identify during these formative years due to their requirement for lifelong care and monitoring. The dataset, collected from 2019 to 2022 by Serbia’s Healthcare Center through an observational cross-sectional study, posed challenges common to medical datasets, including imbalances, data scarcity, and a need for transparent, explainable predictive models. To counter these issues, we utilized three AutoML frameworks - AutoGluon, H2O, and MLJAR - in conjunction with a Tabular Variational Autoencoder (TVAE) to synthetically augment the data points, Prinicipal Component Analysis (PCA) for dimensionality reduction, and SHapley Additive exPlanations (SHAP) and Permutation feature importance analyses to extract insights from the results. AutoGluon outperformed the others on the original dataset, delivering better results with weighted ensemble models for both conditions under a 12-minute budget-time constraint and maintaining all evaluation metrics below a 4% threshold, all without the need for further scaling or calibration in the experimental setup. Our research underscores the broad applicability of the current AutoML paradigm, highlighting its particular benefits for the healthcare domain and diagnostics, where such advanced tools can enhance patient care.