<p>Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model’s stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption—that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss’ <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\kappa\)</EquationSource> </InlineEquation> 0.087–0.223; disagreement 62.2%–74.3%), so the same rationale can be judged supported or unsupported depending on the verifier. A preliminary validation in which a physician adjudicated 50 cases echoed this: agreement with the physician varied across verifiers, underscoring that no single LLM reliably stands in for clinical assessment. Together, these results suggest a single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential. The unanimous-agreement tier offers a candidate for selective automation, but its clinical reliability remains to be confirmed in larger, multi-clinician adjudication studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-Axial Analysis of Clinical Reasoning in Large Language Models: Inter-Verifier Disagreement and Its Implications for Automated Evaluation

  • Hyunjung Byun,
  • Dahyoun Lee,
  • Munyoung Jung,
  • Beakcheol Jang

摘要

Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model’s stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption—that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss’ \(\kappa\) 0.087–0.223; disagreement 62.2%–74.3%), so the same rationale can be judged supported or unsupported depending on the verifier. A preliminary validation in which a physician adjudicated 50 cases echoed this: agreement with the physician varied across verifiers, underscoring that no single LLM reliably stands in for clinical assessment. Together, these results suggest a single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential. The unanimous-agreement tier offers a candidate for selective automation, but its clinical reliability remains to be confirmed in larger, multi-clinician adjudication studies.