Objective <p>To evaluate and compare the performance&#xa0;(or accuracy) of publicly available large language models (ChatGPT-4.0, ChatGPT-3.5, and Google Bard) in answering multiple-choice postgraduate-level surgical examination questions.</p> Methods <p>A search was conducted on PubMed/MEDLINE and the Cochrane Library for studies that compared the accuracy of ChatGPT and Google Bard in the context of multiple-choice postgraduate-level surgical examination questions. A random-effects model was used for statistical analysis to estimate and compare the pooled accuracy&#xa0;of the large language models, with results reported as a 95% confidence interval (CI). Heterogeneity was assessed using the I<sup>2</sup> statistic and publication bias was evaluated through funnel plots and Egger’s test. Statistical significance&#xa0;was set at P &lt; 0.05.</p> Results <p>The full text of 12 studies published between 2023 and 2024 was reviewed, and data extraction was conducted to compare the performance of ChatGPT (GPT-3.5 or GPT-4.0) with Google Bard (rebranded as Gemini). ChatGPT-4.0 exhibited the highest accuracy, with a pooled accuracy of 73% (95% CI: 0.65—0.80, P &lt; 0.01, <i>I</i><sup>2</sup> = 94%). No statistically significant difference was observed when comparing ChatGPT-3.5 with Google Bard (OR: 0.98, 95% CI: 0.8—1.21, P = 0.88, <i>I</i><sup>2</sup> = 67%) which showed that both models&#xa0;performed at a similar level. A statistically significant difference was found when comparing ChatGPT-4.0 with Google Bard (OR: 2.25, 95% CI: 1.73—2.91, P &lt; 0.01, <i>I</i><sup>2</sup> = 78%), showing that ChatGPT-4.0 demonstrated superior performance.</p> Conclusion <p>This meta-analysis highlighted the strong potential of large language models to pass postgraduate-level surgical examinations. Of the three large language models, ChatGPT-4.0 demonstrated the best accuracy, while Google Bard showed the most inconsistent performance, scoring under 50% in 4 of the 12 studies analyzed. These findings suggested that large language models, particularly ChatGPT-4.0, could apply surgical knowledge to solve problems, with the potential for future applications in medical education and patient care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of ChatGPT versus Google Bard on Answering Postgraduate-Level Surgical Examination Questions: A Meta-Analysis

  • Albert Andrew,
  • Sunny Zhao

摘要

Objective

To evaluate and compare the performance (or accuracy) of publicly available large language models (ChatGPT-4.0, ChatGPT-3.5, and Google Bard) in answering multiple-choice postgraduate-level surgical examination questions.

Methods

A search was conducted on PubMed/MEDLINE and the Cochrane Library for studies that compared the accuracy of ChatGPT and Google Bard in the context of multiple-choice postgraduate-level surgical examination questions. A random-effects model was used for statistical analysis to estimate and compare the pooled accuracy of the large language models, with results reported as a 95% confidence interval (CI). Heterogeneity was assessed using the I2 statistic and publication bias was evaluated through funnel plots and Egger’s test. Statistical significance was set at P < 0.05.

Results

The full text of 12 studies published between 2023 and 2024 was reviewed, and data extraction was conducted to compare the performance of ChatGPT (GPT-3.5 or GPT-4.0) with Google Bard (rebranded as Gemini). ChatGPT-4.0 exhibited the highest accuracy, with a pooled accuracy of 73% (95% CI: 0.65—0.80, P < 0.01, I2 = 94%). No statistically significant difference was observed when comparing ChatGPT-3.5 with Google Bard (OR: 0.98, 95% CI: 0.8—1.21, P = 0.88, I2 = 67%) which showed that both models performed at a similar level. A statistically significant difference was found when comparing ChatGPT-4.0 with Google Bard (OR: 2.25, 95% CI: 1.73—2.91, P < 0.01, I2 = 78%), showing that ChatGPT-4.0 demonstrated superior performance.

Conclusion

This meta-analysis highlighted the strong potential of large language models to pass postgraduate-level surgical examinations. Of the three large language models, ChatGPT-4.0 demonstrated the best accuracy, while Google Bard showed the most inconsistent performance, scoring under 50% in 4 of the 12 studies analyzed. These findings suggested that large language models, particularly ChatGPT-4.0, could apply surgical knowledge to solve problems, with the potential for future applications in medical education and patient care.