Optimizing AI in Medical Education: Cost-Benefit Analysis of Large Language Models in the MIR Examination
摘要
This study explores the cost-effectiveness and sustainability of large language models (LLMs) in medical education, focusing on their application in Spain’s MIR examination. It evaluates trade-offs between accuracy, computational cost, and practical feasibility. Findings reveal that Miri Pro, a domain-specific model, surpassed generalist LLMs in both accuracy (195/210) and cost-efficiency, outperforming the best human score. High-end models like GPT-4 Turbo demonstrated advanced reasoning but incurred high costs per correct response. The study highlights the need for cost-optimized, fine-tuned AI solutions and questions the scalability of premium models. It introduces a novel cost-benefit framework, advancing debate on AI’s viability in resource-constrained settings. Practical implications include prioritizing AI literacy, regulatory equity, and financially sustainable AI integration in medical education.