<p>Adolescent major depressive disorder (MDD) is prevalent yet under-recognized, underscoring the need for scalable and objective assessment tools. We developed and validated smartphone-based speech models for detecting MDD in a multi-center cohort of Chinese adolescents (<i>N</i> = 1838) with clinician-established MDD (<i>N</i> = 981), and an external validation cohort (<i>N</i> = 145). A fine-tuned Whisper-based model (Whisper-FT) outperformed an interpretable benchmark using handcrafted acoustic features (ComParE-OS-CB) in development cohort (AUROC 0.95 vs 0.84) and external validation (0.88 vs. 0.82). Performance was stable across age, sex, and symptom-severity subgroups, without significant fairness disparities. For interpretability and generalizability, we further conducted leave-one-site-out validation. Shapley additive explanations (SHAP) analyses of the handcrafted benchmark suggested that spectral-shape, cepstral, and jitter-related patterns drove model discrimination. Decision-curve analyses supported clinical net benefit, and speech showed incremental value beyond PHQ-8. These findings support smartphone speech as a complementary tool for adolescent depression screening. Trial registration: ChiCTR2600116692, registered on 14 January 2026.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An objective smartphone speech diagnostic aid for adolescent major depressive disorder

  • Jiayi Yan,
  • Biman Najika Liyanage,
  • Kangfuxi Zhang,
  • Simin Kang,
  • Zhengwen Zhu,
  • Zaixu Cui,
  • Yu Bai,
  • Yutao Sun,
  • Wenwu Li,
  • Yun Mo,
  • Xiaoyan Xue,
  • Haiyan Zhang,
  • Yulong Li,
  • Ling Jiang,
  • Rui Ren,
  • Jun Yang,
  • Zongfeng Li,
  • Qingjiu Cao

摘要

Adolescent major depressive disorder (MDD) is prevalent yet under-recognized, underscoring the need for scalable and objective assessment tools. We developed and validated smartphone-based speech models for detecting MDD in a multi-center cohort of Chinese adolescents (N = 1838) with clinician-established MDD (N = 981), and an external validation cohort (N = 145). A fine-tuned Whisper-based model (Whisper-FT) outperformed an interpretable benchmark using handcrafted acoustic features (ComParE-OS-CB) in development cohort (AUROC 0.95 vs 0.84) and external validation (0.88 vs. 0.82). Performance was stable across age, sex, and symptom-severity subgroups, without significant fairness disparities. For interpretability and generalizability, we further conducted leave-one-site-out validation. Shapley additive explanations (SHAP) analyses of the handcrafted benchmark suggested that spectral-shape, cepstral, and jitter-related patterns drove model discrimination. Decision-curve analyses supported clinical net benefit, and speech showed incremental value beyond PHQ-8. These findings support smartphone speech as a complementary tool for adolescent depression screening. Trial registration: ChiCTR2600116692, registered on 14 January 2026.