<p>Access to subspecialty oculoplastic triage is limited; multimodal large language model tools may help if accurate, well-calibrated, and clinically useful. This prospective validation study at a tertiary oculoplastics clinic (March–August 2025) evaluated a GPT-4o-based chatbot for diagnostic accuracy, misclassification, calibration, decision-curve utility, and explanation quality. Consecutive adults with eyelid abnormalities were imaged with standardized frontal photographs, and two oculoplastic surgeons established reference diagnoses. Among 261 images (mean age 46.7 years; 72.4% women), overall micro-precision was 0.882, micro-recall 0.839, and micro-F1 0.860. Common diagnoses (ptosis, dermatochalasis) showed high sensitivity and specificity, while rarer classes showed high precision but lower sensitivity. Eyelid tumors had precision 0.96 and recall 0.68, with ptosis, dermatochalasis, and ectropion accounting for 75% of false-negative tumors. Calibration revealed under-calling for tumors (calibration ratio 0.70) and xanthelasma (0.25), and over-calling for lower-lid retraction (1.40). Decision-curve analysis favored the chatbot over refer-all or refer-none strategies across most thresholds. Explanation quality was high (mean 13.3/16) and associated with fewer tumor misses (odds ratio 0.084 per 1-point increase). Multimodal chatbots may support eyelid triage as a proof-of-concept tool, but calibration-aware thresholds and mandatory secondary review for suspected tumors are recommended before deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal chatbot triage for eyelid disorders shows high accuracy and net benefit

  • Hossein Ghahvehchian,
  • Kimia Daneshvar,
  • Amir Manavishad,
  • Hanieh Hosseini,
  • Nasser Karimi,
  • Reza Mirshahi,
  • Zahra Ghomi,
  • Mohsen Bahmani Kashkouli

摘要

Access to subspecialty oculoplastic triage is limited; multimodal large language model tools may help if accurate, well-calibrated, and clinically useful. This prospective validation study at a tertiary oculoplastics clinic (March–August 2025) evaluated a GPT-4o-based chatbot for diagnostic accuracy, misclassification, calibration, decision-curve utility, and explanation quality. Consecutive adults with eyelid abnormalities were imaged with standardized frontal photographs, and two oculoplastic surgeons established reference diagnoses. Among 261 images (mean age 46.7 years; 72.4% women), overall micro-precision was 0.882, micro-recall 0.839, and micro-F1 0.860. Common diagnoses (ptosis, dermatochalasis) showed high sensitivity and specificity, while rarer classes showed high precision but lower sensitivity. Eyelid tumors had precision 0.96 and recall 0.68, with ptosis, dermatochalasis, and ectropion accounting for 75% of false-negative tumors. Calibration revealed under-calling for tumors (calibration ratio 0.70) and xanthelasma (0.25), and over-calling for lower-lid retraction (1.40). Decision-curve analysis favored the chatbot over refer-all or refer-none strategies across most thresholds. Explanation quality was high (mean 13.3/16) and associated with fewer tumor misses (odds ratio 0.084 per 1-point increase). Multimodal chatbots may support eyelid triage as a proof-of-concept tool, but calibration-aware thresholds and mandatory secondary review for suspected tumors are recommended before deployment.