错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams

  • Rodrigo M. Carrillo-Larco

摘要

LLMs have demonstrated remarkable ability in answering medical examinations. However, whether their performance remains stable under stress evaluations is unknown. We used PeruMedQA (n=8,380), a multiple-choice question-answering dataset, and ten medical LLMs. The stress test consisted of randomly shuffling the multiple-choice answers. Using paired t-tests and Wilcoxon tests, we compared LLM accuracy on the original versus the shuffled exams. MedGemma 27B, OctoMed-7B, and Meditron 7B did not exhibit statistically significant differences, overall and stratified by year/specialty. The largest non-significant differences for these LLMs were −2.82, −3.15, and −5.96 percentage points, respectively. These three LLMs may represent robust options for AI applications in Spanish-speaking Latin America.