Aim <p>This study aims to comparatively evaluate the performance of four different artificial intelligence-based chatbots (ChatGPT-4o (Free), ChatGPT-5 (Plus), DeepSeek, and Google Gemini) in the diagnosis and treatment processes of dental trauma cases.</p> Material and methods <p>Based on the International Association of Dental Traumatology (IADT) guidelines, eleven fictional cases were developed, each representing different diagnostic possibilities of dental trauma. The cases cenarios were based on a standardized dataset including clinical examination, radiographic findings, and pulp sensitivity tests. All AI models were tested for three days for each case. Two researchers analyzed the responses using ablinding method according to five evaluation criteria (diagnostic accuracy, treatment plan appropriateness, splinting duration accuracy, antibiotic indication, and reference accuracy).</p> Results <p>According to the results, while no significant difference was found in terms of diagnostic accuracy (<i>p</i> &gt; 0.05), Google Gemini showed the highest performance with 100% accuracy. ChatGPT-4o (Free) stood out with a 97% accuracy rate in antibiotic indication, while in splinting duration prediction, ChatGPT-5 (Plus) was the most successful model with 75.8%. DeepSeek exhibited the highest variability (<i>p</i> &lt; 0.05).</p> Conclusions <p>The findings show that AI chatbots have promising potential as complementary tools in dental trauma diagnosis, but evidence-based validation, expert supervision, and methodological standards are needed for safe use in clinical practice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of artificial intelligence chatbots in the diagnosis and management of simulated dental trauma cases: an evaluation based on IADT guidelines

  • Öznur Küçük Keleş,
  • Zeynep Betül Arslan

摘要

Aim

This study aims to comparatively evaluate the performance of four different artificial intelligence-based chatbots (ChatGPT-4o (Free), ChatGPT-5 (Plus), DeepSeek, and Google Gemini) in the diagnosis and treatment processes of dental trauma cases.

Material and methods

Based on the International Association of Dental Traumatology (IADT) guidelines, eleven fictional cases were developed, each representing different diagnostic possibilities of dental trauma. The cases cenarios were based on a standardized dataset including clinical examination, radiographic findings, and pulp sensitivity tests. All AI models were tested for three days for each case. Two researchers analyzed the responses using ablinding method according to five evaluation criteria (diagnostic accuracy, treatment plan appropriateness, splinting duration accuracy, antibiotic indication, and reference accuracy).

Results

According to the results, while no significant difference was found in terms of diagnostic accuracy (p > 0.05), Google Gemini showed the highest performance with 100% accuracy. ChatGPT-4o (Free) stood out with a 97% accuracy rate in antibiotic indication, while in splinting duration prediction, ChatGPT-5 (Plus) was the most successful model with 75.8%. DeepSeek exhibited the highest variability (p < 0.05).

Conclusions

The findings show that AI chatbots have promising potential as complementary tools in dental trauma diagnosis, but evidence-based validation, expert supervision, and methodological standards are needed for safe use in clinical practice.