<p>With the growing use of language models (LMs) in clinical environments, there is an immediate need to evaluate the accuracy of LMs. Detecting errors in LM-generated text is challenging because (1) manual review is costly and (2) expert-composed reference outputs are often unavailable in real-world settings. While the “LLM-as-a-judge” paradigm (a LM evaluating another LM) offers scalable evaluation, even frontier LMs can miss subtle but clinically significant errors. To address these challenges, we propose MedVAL, a novel, self-supervised, data-efficient distillation method that leverages synthetic data to train evaluator LMs to assess whether LM-generated medical outputs are factually consistent with inputs, without requiring physician labels or reference outputs. To evaluate LM performance, we introduce MedVAL-Bench, a dataset of 840 physician-annotated outputs across 6 diverse clinical use cases capturing real-world challenges. Each output is reviewed following a physician-defined taxonomy of risk levels and error categories, enabling evaluation of LMs in making deployment safety decisions. Across 10 state-of-the-art LMs spanning open-source, proprietary, and medically adapted models, MedVAL distillation significantly improves (<i>p</i> &lt; 0.001) alignment with physicians across seen and unseen tasks, increasing average F1 scores from 66% to 83%. Despite strong baseline performance, MedVAL improves the best-performing proprietary LM (GPT-4o) by 8% without training on physician-labeled data, demonstrating a performance statistically non-inferior to a single human expert on a subset annotated by multiple physicians (<i>p</i> &lt; 0.001). To support a scalable pathway towards clinical integration, we open-source: (1) <a href="https://github.com/StanfordMIMI/MedVAL">Codebase</a>, (2) <a href="https://huggingface.co/datasets/stanfordmimi/MedVAL-Bench">MedVAL-Bench</a>, (3) <a href="https://huggingface.co/stanfordmimi/MedVAL-4B">MedVAL-4B</a>. Across evaluated tasks, our study provides evidence of LMs approaching expert-level ability in risk-aware validation of LM-generated medical text.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward expert-level medical text validation with language models

  • Asad Aali,
  • Vasiliki Bikia,
  • Maya Varma,
  • Nicole Chiou,
  • Sophie Ostmeier,
  • Arnav Singhvi,
  • Magdalini Paschali,
  • Ashwin Kumar,
  • Andrew Johnston,
  • Karimar Amador-Martinez,
  • Eduardo Juan Perez Guerrero,
  • Paola Naovi Cruz Rivera,
  • Sergios Gatidis,
  • Christian Bluethgen,
  • Eduardo Pontes Reis,
  • Eddy D. Zandee van Rilland,
  • Poonam Laxmappa Hosamani,
  • Kevin R. Keet,
  • Minjoung Go,
  • Evelyn Ling,
  • David B. Larson,
  • Curtis Langlotz,
  • Roxana Daneshjou,
  • Jason Hom,
  • Sanmi Koyejo,
  • Emily Alsentzer,
  • Akshay S. Chaudhari

摘要

With the growing use of language models (LMs) in clinical environments, there is an immediate need to evaluate the accuracy of LMs. Detecting errors in LM-generated text is challenging because (1) manual review is costly and (2) expert-composed reference outputs are often unavailable in real-world settings. While the “LLM-as-a-judge” paradigm (a LM evaluating another LM) offers scalable evaluation, even frontier LMs can miss subtle but clinically significant errors. To address these challenges, we propose MedVAL, a novel, self-supervised, data-efficient distillation method that leverages synthetic data to train evaluator LMs to assess whether LM-generated medical outputs are factually consistent with inputs, without requiring physician labels or reference outputs. To evaluate LM performance, we introduce MedVAL-Bench, a dataset of 840 physician-annotated outputs across 6 diverse clinical use cases capturing real-world challenges. Each output is reviewed following a physician-defined taxonomy of risk levels and error categories, enabling evaluation of LMs in making deployment safety decisions. Across 10 state-of-the-art LMs spanning open-source, proprietary, and medically adapted models, MedVAL distillation significantly improves (p < 0.001) alignment with physicians across seen and unseen tasks, increasing average F1 scores from 66% to 83%. Despite strong baseline performance, MedVAL improves the best-performing proprietary LM (GPT-4o) by 8% without training on physician-labeled data, demonstrating a performance statistically non-inferior to a single human expert on a subset annotated by multiple physicians (p < 0.001). To support a scalable pathway towards clinical integration, we open-source: (1) Codebase, (2) MedVAL-Bench, (3) MedVAL-4B. Across evaluated tasks, our study provides evidence of LMs approaching expert-level ability in risk-aware validation of LM-generated medical text.