Comparative performance of large language models in answering periodontology questions from the Turkish Dental Specialty Examination: a cross-sectional study on accuracy and coverage
摘要
In recent years, several studies have explored the use of large language models (LLMs) such as ChatGPT-4, Claude, Gemini Advanced, and DeepSeek-R1 in dental education. Nevertheless, no study has yet reported a comparative evaluation of multiple LLMs specifically on the periodontology section of the Turkish Dental Specialty Examination (DUS), nor analyzed how their performance differs between fundamental knowledge questions and clinical decision-making questions. This study aims to fill this gap by comparing the accuracy and coverage performance of four contemporary LLMs on publicly available DUS periodontology questions.
MethodsA total of 60 publicly available periodontology questions from the DUS (2010–2021) were included. Questions were categorized into Basic Sciences & Pathology and Clinical Applications & Treatment. Each question was administered in its original multiple-choice format (A–E) to ChatGPT-4, Claude, Gemini Advanced, and DeepSeek-R1. Model responses were scored for accuracy (correct/incorrect) and coverage (1–5 rubric). Two independent evaluators assessed coverage, with excellent inter-rater reliability (κ = 0.88). Accuracy rates were compared using Cochran’s Q and McNemar tests, while coverage scores were compared using the Wilcoxon test with Bonferroni correction.
ResultsChatGPT-4 achieved the highest overall accuracy (73.3%), followed by DeepSeek-R1 (63.3%), Gemini Advanced (55.0%), and Claude (36.7%). Accuracy was significantly higher for knowledge-based questions (ChatGPT-4: 80.0%) than for clinical questions (ChatGPT-4: 66.7%). Claude showed the lowest performance in both categories (43.3% and 30.0%). Coverage scores were relatively high across models (means 3.8–4.2) with no statistically significant differences.
ConclusionAmong the tested LLMs, ChatGPT-4 consistently outperformed others in accuracy, while DeepSeek-R1 and Gemini demonstrated moderate performance and Claude lagged behind. Accuracy was lower in clinical questions, reflecting the contextual complexity of clinical reasoning. Coverage scores did not differ significantly, indicating broadly similar comprehensiveness of responses.