<p>The Coronary Artery Disease Reporting and Data System (CAD-RADS) standardizes coronary CT angiography (CCTA) reporting, but not all reports contain CAD-RADS classifications. We benchmarked 54 large language model (LLM) configurations across 50 distinct models, including recent proprietary and open-weight reasoning models, for zero-shot CAD-RADS classification. We retrospectively analyzed 500 anonymized CCTA reports from four hospitals across three U.S. regions. Expert cardiovascular radiologists provided the reference standard (human inter-rater <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\kappa =0.89\)</EquationSource> </InlineEquation>). Fifty-four model configurations (50 distinct LLMs; four configurable models tested in both thinking and non-thinking modes) spanning Llama 2 7B through recent thinking models (DeepCogito v2, Gemini 3 Pro) processed reports using identical zero-shot prompts. Performance was measured with unweighted Cohen’s <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\kappa\)</EquationSource> </InlineEquation>. Two LLMs met both pre-specified non-inferiority criteria (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\delta =0.10\)</EquationSource> </InlineEquation> margin and entire 95% CI within the human inter-rater agreement band, <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\kappa =0.804\)</EquationSource> </InlineEquation>–0.956): Claude 4.6 Opus (<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\kappa =0.853\)</EquationSource> </InlineEquation>, 95% CI 0.817–0.889) and the open-weight Gemma 4 31B (<InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(\kappa =0.845\)</EquationSource> </InlineEquation>, 95% CI 0.810–0.882), which ranked second overall. Both met the criteria at the pre-specified <InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(\delta =0.10\)</EquationSource> </InlineEquation> margin; at <InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\delta =0.05\)</EquationSource> </InlineEquation> no model qualified. On the reports originally dictated without a CAD-RADS statement (n=343), the top models reached <InlineEquation ID="IEq9"> <EquationSource Format="TEX">\(\kappa \approx 0.80\)</EquationSource> </InlineEquation>. Performance declined with longer thinking chains (proxy for case complexity), but thinking-mode outperformed non-thinking mode on matched difficult reports. Gemma 4 31B remained non-inferior at 3-bit quantization and fits on a 24&#xa0;GB consumer GPU. Current LLMs extract the CAD-RADS stenosis severity category from unstructured CCTA reports with agreement approaching the human inter-rater band, without task-specific training. Our data shows the remarkable rise of open models. A 31B open-weight model matched top proprietary systems and runs on consumer GPUs, enabling privacy-preserving local deployment for clinical data mining.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking 54 large language model configurations for CAD-RADS scoring: open-weight models approach human-level agreement

  • Veit Sandfort,
  • Davis M. Vigneault,
  • Martin J. Willemink,
  • Jie Wu,
  • Richard L. Hallett,
  • Koen Nieman,
  • Dominik Fleischmann,
  • Domenico Mastrodicasa

摘要

The Coronary Artery Disease Reporting and Data System (CAD-RADS) standardizes coronary CT angiography (CCTA) reporting, but not all reports contain CAD-RADS classifications. We benchmarked 54 large language model (LLM) configurations across 50 distinct models, including recent proprietary and open-weight reasoning models, for zero-shot CAD-RADS classification. We retrospectively analyzed 500 anonymized CCTA reports from four hospitals across three U.S. regions. Expert cardiovascular radiologists provided the reference standard (human inter-rater \(\kappa =0.89\) ). Fifty-four model configurations (50 distinct LLMs; four configurable models tested in both thinking and non-thinking modes) spanning Llama 2 7B through recent thinking models (DeepCogito v2, Gemini 3 Pro) processed reports using identical zero-shot prompts. Performance was measured with unweighted Cohen’s \(\kappa\) . Two LLMs met both pre-specified non-inferiority criteria ( \(\delta =0.10\) margin and entire 95% CI within the human inter-rater agreement band, \(\kappa =0.804\) –0.956): Claude 4.6 Opus ( \(\kappa =0.853\) , 95% CI 0.817–0.889) and the open-weight Gemma 4 31B ( \(\kappa =0.845\) , 95% CI 0.810–0.882), which ranked second overall. Both met the criteria at the pre-specified \(\delta =0.10\) margin; at \(\delta =0.05\) no model qualified. On the reports originally dictated without a CAD-RADS statement (n=343), the top models reached \(\kappa \approx 0.80\) . Performance declined with longer thinking chains (proxy for case complexity), but thinking-mode outperformed non-thinking mode on matched difficult reports. Gemma 4 31B remained non-inferior at 3-bit quantization and fits on a 24 GB consumer GPU. Current LLMs extract the CAD-RADS stenosis severity category from unstructured CCTA reports with agreement approaching the human inter-rater band, without task-specific training. Our data shows the remarkable rise of open models. A 31B open-weight model matched top proprietary systems and runs on consumer GPUs, enabling privacy-preserving local deployment for clinical data mining.

Graphical Abstract