PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams
摘要
LLMs have demonstrated remarkable ability in answering medical examinations. However, whether their performance remains stable under stress evaluations is unknown. We used PeruMedQA (n=8,380), a multiple-choice question-answering dataset, and ten medical LLMs. The stress test consisted of randomly shuffling the multiple-choice answers. Using paired t-tests and Wilcoxon tests, we compared LLM accuracy on the original versus the shuffled exams. MedGemma 27B, OctoMed-7B, and Meditron 7B did not exhibit statistically significant differences, overall and stratified by year/specialty. The largest non-significant differences for these LLMs were −2.82, −3.15, and −5.96 percentage points, respectively. These three LLMs may represent robust options for AI applications in Spanish-speaking Latin America.