Background <p>Large language models (LLMs), particularly successive iterations of ChatGPT (GPT-3.5 and GPT-4), have demonstrated increasingly sophisticated performance on standardised medical assessments and are now widely accessible to undergraduate medical students worldwide. Although the evidence base has expanded rapidly since ChatGPT’s public release in November 2022, findings remain fragmented across three largely distinct research domains: examination performance, knowledge acquisition, and learner perceptions.</p> Methods <p>A PRISMA 2020-compliant systematic review was conducted. Structured searches were performed in PubMed (MEDLINE), Scopus, and Web of Science Core Collection (January 2022–March 2026), supplemented by grey literature and citation tracking. Eligible studies examined ChatGPT (GPT-3.5, GPT-4, GPT-4o) or comparable transformer-based LLMs in undergraduate medical educational contexts. Two reviewers independently screened records, extracted data, and assessed risk of bias (Newcastle-Ottawa Scale; ROBINS-I). Narrative synthesis followed the Synthesis Without Meta-analysis (SWiM) reporting guideline.</p> Results <p>Of 1,847 records identified, 43 studies (approximately 12,400 participants across 19 countries) met inclusion criteria. GPT-4 achieved passing scores on the United States Medical Licensing Examination (USMLE) Steps 1–3 in all studies in which it was tested (mean accuracy 86%; IQR 81%--89%), significantly exceeding GPT-3.5 performance (mean 56%). Accuracy declined on multi-step clinical vignettes (61%–71%) and image-dependent items (48%–63%). Six of twelve studies reported modest short-term knowledge acquisition benefits effect sizes ranged from 0.31 to 0.71 (median⁓0.53) across six studies without durable retention. Students broadly endorsed LLMs as educational aids (74% positive perception); however, 61% raised academic integrity concerns, 59% cited hallucination risk, and 82% reported receiving no institutional guidance. Faculty expressed more tempered views (51% positive). Risk of bias was moderate-to-high in 86% of included studies.</p> Conclusions <p>GPT-4 reaches examination-passing thresholds on major medical licensing assessments but demonstrates clinically important limitations in multi-step reasoning, uncertainty calibration, and multimodal task performance. The near-universal absence of institutional guidance represents the most urgent and actionable finding. Longitudinal controlled research with validated outcome measures and development of standardised AI literacy curricula are the field’s most pressing priorities.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large language models in undergraduate medical education: a systematic review of examination performance, knowledge acquisition, and learner perceptions

  • Parag Bodke,
  • Rahul Jumle

摘要

Background

Large language models (LLMs), particularly successive iterations of ChatGPT (GPT-3.5 and GPT-4), have demonstrated increasingly sophisticated performance on standardised medical assessments and are now widely accessible to undergraduate medical students worldwide. Although the evidence base has expanded rapidly since ChatGPT’s public release in November 2022, findings remain fragmented across three largely distinct research domains: examination performance, knowledge acquisition, and learner perceptions.

Methods

A PRISMA 2020-compliant systematic review was conducted. Structured searches were performed in PubMed (MEDLINE), Scopus, and Web of Science Core Collection (January 2022–March 2026), supplemented by grey literature and citation tracking. Eligible studies examined ChatGPT (GPT-3.5, GPT-4, GPT-4o) or comparable transformer-based LLMs in undergraduate medical educational contexts. Two reviewers independently screened records, extracted data, and assessed risk of bias (Newcastle-Ottawa Scale; ROBINS-I). Narrative synthesis followed the Synthesis Without Meta-analysis (SWiM) reporting guideline.

Results

Of 1,847 records identified, 43 studies (approximately 12,400 participants across 19 countries) met inclusion criteria. GPT-4 achieved passing scores on the United States Medical Licensing Examination (USMLE) Steps 1–3 in all studies in which it was tested (mean accuracy 86%; IQR 81%--89%), significantly exceeding GPT-3.5 performance (mean 56%). Accuracy declined on multi-step clinical vignettes (61%–71%) and image-dependent items (48%–63%). Six of twelve studies reported modest short-term knowledge acquisition benefits effect sizes ranged from 0.31 to 0.71 (median⁓0.53) across six studies without durable retention. Students broadly endorsed LLMs as educational aids (74% positive perception); however, 61% raised academic integrity concerns, 59% cited hallucination risk, and 82% reported receiving no institutional guidance. Faculty expressed more tempered views (51% positive). Risk of bias was moderate-to-high in 86% of included studies.

Conclusions

GPT-4 reaches examination-passing thresholds on major medical licensing assessments but demonstrates clinically important limitations in multi-step reasoning, uncertainty calibration, and multimodal task performance. The near-universal absence of institutional guidance represents the most urgent and actionable finding. Longitudinal controlled research with validated outcome measures and development of standardised AI literacy curricula are the field’s most pressing priorities.