<p>This study aims to compare and evaluate the accuracy, consistency, and response behavior of ChatGPT-4o and DeepSeek in orthodontic health literacy tasks. A total of 50 multiple-choice, text-based questions covering five orthodontic domains were developed using the Delphi method by a panel of five experienced orthodontic professionals. Each question was asked to both models three times daily (morning, noon, evening) over three days, in both English and Chinese versions. Accuracy, consistency, and response time were analyzed using SPSS and R software. The overall accuracy of ChatGPT-4o (90.4%) was slightly higher than that of DeepSeek (88.0%), though the difference was not statistically significant in either English (<i>p</i> = 1.000) or Chinese (<i>p</i> = 0.263). Across individual question groups, no significant differences in accuracy were found between models (<i>p</i> &gt; 0.05). The time of day and testing day had no effect on model performance (<i>p</i> &gt; 0.05 for all comparisons). Both models demonstrated similar consistency in their responses (<i>p</i> = 1.000 for English; <i>p</i> = 0.412 for Chinese). However, ChatGPT-4o showed significantly faster response times than DeepSeek in both languages (<i>p</i> &lt; 0.01). Both ChatGPT-4o and DeepSeek demonstrate comparable accuracy and consistency in orthodontic health literacy tasks. However, ChatGPT-4o’s significantly faster response time makes it more suitable for clinical applications requiring efficient information delivery.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of ChatGPT-4o and DeepSeek as tools for orthodontic health literacy in public dental education

  • Zhaoxiang Wen,
  • Jiaxin Huang,
  • Keer Yu,
  • Yaqi Li,
  • Zhenhui Wang,
  • Xiaozhu Liao,
  • Biao Li,
  • Zhendong Tao,
  • Hong He

摘要

This study aims to compare and evaluate the accuracy, consistency, and response behavior of ChatGPT-4o and DeepSeek in orthodontic health literacy tasks. A total of 50 multiple-choice, text-based questions covering five orthodontic domains were developed using the Delphi method by a panel of five experienced orthodontic professionals. Each question was asked to both models three times daily (morning, noon, evening) over three days, in both English and Chinese versions. Accuracy, consistency, and response time were analyzed using SPSS and R software. The overall accuracy of ChatGPT-4o (90.4%) was slightly higher than that of DeepSeek (88.0%), though the difference was not statistically significant in either English (p = 1.000) or Chinese (p = 0.263). Across individual question groups, no significant differences in accuracy were found between models (p > 0.05). The time of day and testing day had no effect on model performance (p > 0.05 for all comparisons). Both models demonstrated similar consistency in their responses (p = 1.000 for English; p = 0.412 for Chinese). However, ChatGPT-4o showed significantly faster response times than DeepSeek in both languages (p < 0.01). Both ChatGPT-4o and DeepSeek demonstrate comparable accuracy and consistency in orthodontic health literacy tasks. However, ChatGPT-4o’s significantly faster response time makes it more suitable for clinical applications requiring efficient information delivery.