Background <p>Machine learning models for Obstructive Sleep Apnea (OSA) diagnosis have largely inherited some structural limitations: reliance on generic, opportunistically collected feature sets; use of the Apnea-Hypopnea Index (AHI) as the sole ground truth; poor performance in multi-class severity grading; and predictions that offer clinicians no mechanistic insight. This study addresses these gaps by prospectively assembling a multi-domain dataset that, alongside established demographic, anthropometric, and questionnaire-based predictors, incorporates a panel of craniofacial and intraoral metrics specifically designed to capture the structural-anatomical contributors to OSA — integrating these into an interpretable framework for three-class severity classification evaluated against both AHI and the Oxygen Desaturation Index (ODI).</p> Methods <p>In this single-center study, 233 treatment-naïve adults from a tertiary referral cohort (61.8% severe OSA prevalence) underwent in-laboratory polysomnography (PSG). All predictor variables were collected prior to PSG outcome disclosure through a standardized clinical examination, requiring no overnight recording or specialized equipment. An Artificial Neural Network (ANN) was independently trained for three-class severity classification (No/Mild, Moderate, Severe) for each index. Model performance was evaluated on an independent test set (<i>n</i> = 47; 20% of the sample), with interpretability assessed using SHapley Additive exPlanations (SHAP). Comparison with an anatomy-excluded ablation model was conducted to establish the added value of the full feature set.</p> Results <p>The AHI-based model achieved 87.2% overall accuracy (sensitivity/specificity: No/Mild 0.93/0.97, Moderate 0.80/0.91, Severe 0.88/0.93). The ODI-based model achieved 76.6% accuracy, offering reliable exclusion of severe desaturation burden (No/Mild specificity: 0.94). Univariate analyses confirmed significant associations between OSA severity and STOP-BANG score, age, BMI, neck circumference, observed apnea, loud snoring, high blood pressure, Cervico-Mental Angle, Mentocervical Distance, and submental fat (all <i>p</i> ≤ .034 for both indices). SHAP analysis further identified V-shaped maxillary arch, Mallampati score, increased overjet, and alcohol use as influential model predictors — several reaching high model rankings despite modest univariate significance. Notably, AHI and ODI models diverged in their feature weighting — anatomy-driven features dominated AHI prediction while body habitus and comorbidity markers dominated ODI. Against a conventional demographic and questionnaire-based ablation model, the full anatomy-inclusive ANN achieved substantially higher accuracy (87.2% vs. 72.3%), with the largest gain at the Moderate-class boundary (sensitivity: 0.80 vs. 0.58).</p> Conclusions <p>As a proof-of-concept, this study demonstrates that an interpretable ML framework integrating craniofacial and intraoral assessments with standard clinical predictors can classify OSA severity across three classes and provide feature-level explanations to support clinical reasoning. By developing parallel AHI and ODI models, the framework moves beyond AHI-only paradigms, though both remain frequency-based surrogates; hypoxic burden — quantifying the cumulative oxygen desaturation load per sleep period — is the more physiologically complete target toward which this line of work should progress. Findings are limited by single-center design, spectrum bias from a tertiary referral cohort, modest sample size, and absence of inter-rater reliability data. External validation in larger, more representative populations is needed to confirm the robustness and clinical utility of this approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Three-class obstructive sleep apnea severity assessment: a parallel AHI and ODI explainable artificial intelligence framework using craniofacial-enriched clinical data

  • Zahra Ameli Mazandarani,
  • Mohammad Behnaz,
  • Hamed AmiriFard,
  • Asghar Ebadifar,
  • Hoori Mirmohammadsadeghi,
  • Kazem Dalaie,
  • Shahab Kavousinejad

摘要

Background

Machine learning models for Obstructive Sleep Apnea (OSA) diagnosis have largely inherited some structural limitations: reliance on generic, opportunistically collected feature sets; use of the Apnea-Hypopnea Index (AHI) as the sole ground truth; poor performance in multi-class severity grading; and predictions that offer clinicians no mechanistic insight. This study addresses these gaps by prospectively assembling a multi-domain dataset that, alongside established demographic, anthropometric, and questionnaire-based predictors, incorporates a panel of craniofacial and intraoral metrics specifically designed to capture the structural-anatomical contributors to OSA — integrating these into an interpretable framework for three-class severity classification evaluated against both AHI and the Oxygen Desaturation Index (ODI).

Methods

In this single-center study, 233 treatment-naïve adults from a tertiary referral cohort (61.8% severe OSA prevalence) underwent in-laboratory polysomnography (PSG). All predictor variables were collected prior to PSG outcome disclosure through a standardized clinical examination, requiring no overnight recording or specialized equipment. An Artificial Neural Network (ANN) was independently trained for three-class severity classification (No/Mild, Moderate, Severe) for each index. Model performance was evaluated on an independent test set (n = 47; 20% of the sample), with interpretability assessed using SHapley Additive exPlanations (SHAP). Comparison with an anatomy-excluded ablation model was conducted to establish the added value of the full feature set.

Results

The AHI-based model achieved 87.2% overall accuracy (sensitivity/specificity: No/Mild 0.93/0.97, Moderate 0.80/0.91, Severe 0.88/0.93). The ODI-based model achieved 76.6% accuracy, offering reliable exclusion of severe desaturation burden (No/Mild specificity: 0.94). Univariate analyses confirmed significant associations between OSA severity and STOP-BANG score, age, BMI, neck circumference, observed apnea, loud snoring, high blood pressure, Cervico-Mental Angle, Mentocervical Distance, and submental fat (all p ≤ .034 for both indices). SHAP analysis further identified V-shaped maxillary arch, Mallampati score, increased overjet, and alcohol use as influential model predictors — several reaching high model rankings despite modest univariate significance. Notably, AHI and ODI models diverged in their feature weighting — anatomy-driven features dominated AHI prediction while body habitus and comorbidity markers dominated ODI. Against a conventional demographic and questionnaire-based ablation model, the full anatomy-inclusive ANN achieved substantially higher accuracy (87.2% vs. 72.3%), with the largest gain at the Moderate-class boundary (sensitivity: 0.80 vs. 0.58).

Conclusions

As a proof-of-concept, this study demonstrates that an interpretable ML framework integrating craniofacial and intraoral assessments with standard clinical predictors can classify OSA severity across three classes and provide feature-level explanations to support clinical reasoning. By developing parallel AHI and ODI models, the framework moves beyond AHI-only paradigms, though both remain frequency-based surrogates; hypoxic burden — quantifying the cumulative oxygen desaturation load per sleep period — is the more physiologically complete target toward which this line of work should progress. Findings are limited by single-center design, spectrum bias from a tertiary referral cohort, modest sample size, and absence of inter-rater reliability data. External validation in larger, more representative populations is needed to confirm the robustness and clinical utility of this approach.