Comparative evaluation of AI-driven large language models for periodontal–orthodontic patient questions: accuracy, comprehensiveness, readability, and temporal consistency
摘要
Artificial intelligence (AI)-driven large language models (LLMs) are increasingly used to provide patient-oriented oral health information; however, their performance in answering periodontal–orthodontic patient questions remains unclear. This study aimed to compare the performance of contemporary AI-driven LLMs in answering periodontal–orthodontic patient questions by assessing response accuracy, comprehensiveness, readability, and temporal consistency.
MethodsThirty patient-oriented periodontal–orthodontic questions were submitted to six contemporary LLMs (ChatGPT 5.5, Claude Sonnet 4.6, DeepSeek 3.2, Gemini 3.1 Pro, Grok 4.20, and PerioGPT) under standardized conditions, and the generated responses were analyzed. Accuracy and comprehensiveness were evaluated using a modified five-point Likert scale, whereas readability was assessed using the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. Temporal consistency was assessed by repeating the response-generation process one week later. Inter-model comparisons and temporal consistency analyses were performed using repeated-measures ANOVA or Friedman tests with Bonferroni-adjusted post hoc comparisons, as appropriate.
ResultsSignificant differences were observed among the evaluated LLMs in accuracy (p < 0.001, Kendall’s W = 0.401), comprehensiveness (p < 0.001, Kendall’s W = 0.575), FRE (p < 0.001, partial η2 = 0.807), and FKGL (p < 0.001, partial η2 = 0.688). PerioGPT and Gemini achieved the highest accuracy (4.93 ± 0.13 and 4.90 ± 0.19, respectively) and comprehensiveness scores (4.85 ± 0.21 and 4.90 ± 0.22, respectively). PerioGPT generated highly accurate and comprehensive responses but exhibited the lowest FRE (12.7 ± 10.2) and highest FKGL (15.1 ± 2.1) scores, whereas Claude, DeepSeek, and Grok produced more readable outputs. Although some significant differences in temporal consistency were observed among the evaluated models, the associated effect sizes were generally small (0.041–0.143), indicating broadly similar performance over time.
ConclusionsWhile PerioGPT and Gemini achieved the highest accuracy and comprehensiveness scores, and Claude, DeepSeek, and Grok produced more readable responses, the evaluated models generally demonstrated similar temporal consistency over time; however, no model consistently outperformed the others across all evaluated criteria. Therefore, AI-driven LLMs may support patient education and oral health literacy in periodontal–orthodontic care; however, their use as complementary tools under expert supervision is recommended.