Evaluating Large Language Models for Healthcare: Insights from MCQ Evaluation
摘要
This study investigates the performance of general and medical-specific Large Language Models (LLMs) in obstetrics and gynecology, focusing on their ability to accurately handle medical multiple-choice questions (MCQs). We evaluated models like Llama2, Mistral, PMC_LLaMA, and BioMistral, to assess and enhance their reliability and accuracy. Despite the expectations, general-purpose models occasionally outperformed specialized medical models. Our methods, including Structural Influence Testing and Contextual Enhancement Testing, demonstrated significant potential in improving model accuracy and reducing misinformation. Specifically, Structural Influence Testing increased Mistral’s accuracy from 40% to 46% and Llama2’s from 28% to 43% with five shots. Contextual Enhancement Testing yielded a 4% accuracy gain for Mistral and 6% for Llama2 using search terms. This research highlights the importance of optimizing LLMs to empower healthcare professionals with precise and reliable medical information, ultimately improving patient outcomes and supporting informed clinical decisions.