Purpose <p>The aim of this study was to evaluate the performance of GPT-4 in answering oral and maxillofacial surgery (OMFS) board exam questions, given its success in other medical specializations.</p> Methods <p>A total of 250 multiple-choice questions were randomly selected from an established OMFS question bank, covering a broad range of topics such as craniofacial trauma, oncological procedures, orthognathic surgery, and general surgical principles. GPT-4’s responses were assessed for accuracy, and statistical analysis was performed to compare its performance across different topics.</p> Results <p>GPT-4 achieved an overall accuracy of 62% in answering the OMFS board exam questions. The highest accuracies were observed in Pharmacology (92.8%), Anatomy (73.3%), and Mucosal Lesions (70.8%). Conversely, the lowest accuracies were noted in Dental Implants (37.5%), Orthognathic Surgery (38.5%), and Reconstructive Surgery (42.9%). Statistical analysis indicated significant variability in performance across different topics, with GPT-4 performing better in general topics compared to specialized ones.</p> Conclusion <p>GPT-4 demonstrates a promising ability to answer OMFS board exam questions, particularly in general medical topics. However, its performance in highly specialized areas reveals significant limitations. These findings suggest that while GPT-4 can be a useful tool in medical education, further enhancements are needed for its application in specialized medical fields.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of GPT-4 in oral and maxillofacial surgery board exams: challenges in specialized questions

  • Felix Benjamin Warwas,
  • Nils Heim

摘要

Purpose

The aim of this study was to evaluate the performance of GPT-4 in answering oral and maxillofacial surgery (OMFS) board exam questions, given its success in other medical specializations.

Methods

A total of 250 multiple-choice questions were randomly selected from an established OMFS question bank, covering a broad range of topics such as craniofacial trauma, oncological procedures, orthognathic surgery, and general surgical principles. GPT-4’s responses were assessed for accuracy, and statistical analysis was performed to compare its performance across different topics.

Results

GPT-4 achieved an overall accuracy of 62% in answering the OMFS board exam questions. The highest accuracies were observed in Pharmacology (92.8%), Anatomy (73.3%), and Mucosal Lesions (70.8%). Conversely, the lowest accuracies were noted in Dental Implants (37.5%), Orthognathic Surgery (38.5%), and Reconstructive Surgery (42.9%). Statistical analysis indicated significant variability in performance across different topics, with GPT-4 performing better in general topics compared to specialized ones.

Conclusion

GPT-4 demonstrates a promising ability to answer OMFS board exam questions, particularly in general medical topics. However, its performance in highly specialized areas reveals significant limitations. These findings suggest that while GPT-4 can be a useful tool in medical education, further enhancements are needed for its application in specialized medical fields.