Comparative diagnostic accuracy of ChatGPT models in salivary gland disease: a multimodal vignette-based evaluation
摘要
This study evaluated the diagnostic accuracy and consistency of ChatGPT-4o in salivary gland disorders compared to experienced clinicians.
MethodsEighty anonymized salivary gland cases from peer-reviewed reports were evaluated by ChatGPT-4o using standardized multimodal prompts and by three oral medicine specialists who provided Top-5 differentials. The primary outcome was diagnostic accuracy at the most likely diagnosis (Top-1), within the top three (Top-3), and within the top five (Top-5) differential diagnoses, with agreement measured by Cohen’s kappa and subgroup analyses by gland type, imaging, and case difficulty.
ResultsAt Top-3 and Top-5, ChatGPT showed perfect sensitivity (100%) and Top-1 86.67%. Experts surpassed ChatGPT at Top-5 (77.5% vs. 67.5%, p < 0.0001), but ChatGPT outperformed experts at Top-1 (50.0% vs. 37.5%, p = 0.0309) and Top-3 (62.5% vs. 62.5%, p = 1.000). At Top-1, Cohen’s Kappa indicated moderate agreement (0.55). Experts showed notable variation by modality (p = 0.0174) and gland (p = 0.053). Although initial subgroup analyses found no notable heterogeneity of ChatGPT’s performance across imaging modalities, multivariate regression identified gland type to be an independent predictor of its Top-1 accuracy.
ConclusionsThis first study shows ChatGPT can provide expert-level differential diagnoses for salivary gland disorders, suggesting promise as a supportive tool, though further research is needed to confirm its clinical role.
Clinical significanceChatGPT-4o shows promise as a reliable supportive tool for differential diagnosis in oral medicine. Compared to experts, it performed more consistently across imaging modalities, although the particular salivary gland involved had a significant impact on its accuracy. Further validation through larger studies is needed for its integration into routine clinical practice.