<p>This study evaluated the impact of guideline-based prompting on the performance of large language models (LLMs) in answering dental trauma–related questions. Sixteen multiple-choice questions (MCQs) and sixteen open-ended questions (OEQs) derived from the International Association of Dental Traumatology (IADT) guidelines were used. ChatGPT-4o, Gemini-2.5 Flash, and DeepSeek v3.2 were tested under two conditions: with and without guideline support. Questions were asked three times daily over three consecutive days using independent chat sessions. In the guideline-based condition, the dental trauma guideline was uploaded and models were instructed to answer according to the document. Responses were evaluated using predefined answer keys and a structured rubric. Statistical analyses were performed using the Shapiro–Wilk test, aligned rank transform (ART) analysis, and Bonferroni-adjusted multiple comparisons. Multiple-choice performance was consistently high across all models, with no significant effects of day, time, or AI model. Guideline-based prompting substantially improved performance and guideline concordance on open-ended questions. In the guideline-supported condition, DeepSeek achieved a median score of 32, while ChatGPT and Gemini achieved median scores of 30. Without guideline support, median scores decreased to 26, 20, and 23, respectively. All models achieved 100% accuracy on MCQs with guideline support, whereas accuracy ranged from 81.2% to 100% without guideline support. Significant differences between AI models were observed for OEQ scores (<i>p</i> &lt; 0.001). Guideline-based prompting improved guideline concordance and overall performance, particularly for open-ended dental trauma questions. These findings support the use of guideline grounding to enhance the guideline concordance and reliability of LLM-generated responses while emphasizing the continued need for clinician oversight.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Impact of guideline-based prompting on the large language model performance in dental trauma management clinical decision-making

  • Sena Kaşıkçı,
  • Ebru Şirinoğlu,
  • Olcay Özdemir

摘要

This study evaluated the impact of guideline-based prompting on the performance of large language models (LLMs) in answering dental trauma–related questions. Sixteen multiple-choice questions (MCQs) and sixteen open-ended questions (OEQs) derived from the International Association of Dental Traumatology (IADT) guidelines were used. ChatGPT-4o, Gemini-2.5 Flash, and DeepSeek v3.2 were tested under two conditions: with and without guideline support. Questions were asked three times daily over three consecutive days using independent chat sessions. In the guideline-based condition, the dental trauma guideline was uploaded and models were instructed to answer according to the document. Responses were evaluated using predefined answer keys and a structured rubric. Statistical analyses were performed using the Shapiro–Wilk test, aligned rank transform (ART) analysis, and Bonferroni-adjusted multiple comparisons. Multiple-choice performance was consistently high across all models, with no significant effects of day, time, or AI model. Guideline-based prompting substantially improved performance and guideline concordance on open-ended questions. In the guideline-supported condition, DeepSeek achieved a median score of 32, while ChatGPT and Gemini achieved median scores of 30. Without guideline support, median scores decreased to 26, 20, and 23, respectively. All models achieved 100% accuracy on MCQs with guideline support, whereas accuracy ranged from 81.2% to 100% without guideline support. Significant differences between AI models were observed for OEQ scores (p < 0.001). Guideline-based prompting improved guideline concordance and overall performance, particularly for open-ended dental trauma questions. These findings support the use of guideline grounding to enhance the guideline concordance and reliability of LLM-generated responses while emphasizing the continued need for clinician oversight.