错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating large language models on multilingual vaccine knowledge: a benchmark study

  • Siyuan Chen,
  • Lily Wass,
  • Zhengdong Wu,
  • Lucas Garay,
  • José Vizoso,
  • Kathy Leung,
  • Joseph Wu,
  • Leesa Lin

摘要

Large language models (LLMs) are increasingly used by clinicians and the public for vaccine information, yet their factual accuracy across languages and vaccine domains remains insufficiently characterized. We evaluated 13 LLMs using VaxEval, a multilingual vaccine-knowledge benchmark of 1886 vaccine-related multiple-choice questions spanning 14 vaccines in English (71%), Spanish (13%), and Chinese (16%). All items underwent quality control, with reference answers verified against authoritative guidance and peer-reviewed sources. Model performance was evaluated under zero-shot, few-shot, and chain-of-thought (CoT) prompting, with exact-match accuracy defined as selecting the pre-specified reference option. We used mixed-effect logistic regression to estimate associations between model group (newer flagship models vs earlier models), prompting strategy, language, and vaccine type, and answer correctness. Mean accuracy across models was 86.0% in English, 83.7% in Spanish, and 80.0% in Chinese. Flagship models had higher odds of correctness than earlier versions (OR 1.57; 95% CI 1.50–1.65; P < .001). Few-shot prompting was associated with higher correctness (OR 1.17; P < .001), whereas CoT prompting was associated with lower correctness (OR 0.79; P < .001). Performance varied by vaccine type and question category, underscoring the need for rigorous evaluation, structured guardrails, and targeted refinement before using LLMs for vaccine communication.