Concordance between large language models and orthodontists in photo-based orthodontic assessment
摘要
The widespread use of digital platforms for preliminary health information has increased interest in artificial intelligence (AI)-based tools in dentistry. Large language model (LLM)-based systems are increasingly used for patient education and preliminary clinical guidance. However, their clinical consistency in orthodontic assessments based solely on intraoral photographs remains unclear. This study aimed to evaluate the clinical appropriateness of LLM-generated orthodontic assessments compared with blinded expert orthodontist evaluations.
MethodsThis blinded comparative agreement study included 104 patients aged ≥ 12 years undergoing orthodontic evaluation. For each case, three standardized intraoral photographs and basic demographic data were submitted to two LLM-based systems (ChatGPT-4o and Grok 4.1) using a standardized prompt. Responses to five clinical questions were independently evaluated by three expert orthodontists using a 5-point Likert scale. Consensus scores were obtained by averaging the ratings of the three evaluators. Inter-system comparisons were performed using the Wilcoxon signed-rank test. Bonferroni and Benjamini–Hochberg false discovery rate (FDR) corrections were applied for multiple comparisons.
ResultsMean ratings were generally ≥ 3, indicating moderate clinical acceptability. After correction for multiple comparisons, the most robust finding was the significantly higher rating assigned to ChatGPT-4o for the recommended orthodontic treatment approach in Dental Class I patients. In the 12–15-year subgroup, ChatGPT-4o also received higher ratings for treatment recommendation; however, statistical significance was retained only after FDR correction and not after Bonferroni adjustment. No significant differences were observed across most other evaluated parameters, particularly malocclusion classification. In patients older than 15 years, paired differences between the two systems were highly similar, resulting in an insufficient number of non-zero differences for reliable Wilcoxon signed-rank testing.
ConclusionsLLM-based systems may provide clinically acceptable outputs in selected orthodontic assessment domains when evaluations are based solely on intraoral photographs. However, variability persists in complex decision-making tasks, particularly treatment planning. Current evidence supports the use of such systems as supportive tools operating under professional supervision rather than as independent clinical decision-makers.