<p>Extracting disease labels from radiology reports is essential for developing deep learning-based diagnostic models and enabling large-scale retrospective clinical research. Classification of usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports according to Fleischner Society guidelines is a particularly demanding task, requiring synthesis of spatial distribution, fibrotic features, and exclusion criteria. As open-source large language models (LLMs) are released at an accelerating pace with steadily improving general benchmarks, a practical question arises: Do these improvements translate to better performance on complex, real-world clinical classification, and does the optimal prompting strategy differ across model architectures? While prior studies have evaluated LLMs for radiology report labeling, none have compared how prompting strategies interact with the native reasoning capabilities of newer model architectures. We evaluated 10 open-source LLMs from three architecture families (Llama, Qwen, Gemma) spanning 8 to 405 billion parameters, each tested with three prompting strategies on 270 HRCT reports classified by expert consensus of two senior thoracic radiologists. Four models with native reasoning (“thinking”) capability were additionally tested in thinking mode. The best configuration achieved a Cohen’s kappa (<i>κ</i>) of 0.70 and 82% four-class accuracy. Structured reasoning prompting improved all Llama models but degraded all models with native reasoning capability (Qwen 3.5 and Gemma 4), revealing an architecture-dependent interaction. Thinking mode hurt performance on criteria-based prompts and never yielded the best configuration. Larger models did not consistently outperform smaller ones: Llama 3.1 405B offered no advantage over Llama 3.3 70B, and the Qwen 397B model underperformed the dense Qwen 27B. These findings demonstrate that newer model generations with improved general benchmarks and larger parameter counts do not guarantee better performance on specialized medical classification tasks and that prompt design must be matched to model architecture.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large Language Models for Classifying Usual Interstitial Pneumonia from Radiology Reports: Native Reasoning Versus Structured Prompting

  • Ran Zhang,
  • Thomas M. Grist,
  • Mark Schiebler,
  • Yijing Wu,
  • Nathan Sandbo,
  • Allan R. Brasier,
  • Guang-Hong Chen

摘要

Extracting disease labels from radiology reports is essential for developing deep learning-based diagnostic models and enabling large-scale retrospective clinical research. Classification of usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports according to Fleischner Society guidelines is a particularly demanding task, requiring synthesis of spatial distribution, fibrotic features, and exclusion criteria. As open-source large language models (LLMs) are released at an accelerating pace with steadily improving general benchmarks, a practical question arises: Do these improvements translate to better performance on complex, real-world clinical classification, and does the optimal prompting strategy differ across model architectures? While prior studies have evaluated LLMs for radiology report labeling, none have compared how prompting strategies interact with the native reasoning capabilities of newer model architectures. We evaluated 10 open-source LLMs from three architecture families (Llama, Qwen, Gemma) spanning 8 to 405 billion parameters, each tested with three prompting strategies on 270 HRCT reports classified by expert consensus of two senior thoracic radiologists. Four models with native reasoning (“thinking”) capability were additionally tested in thinking mode. The best configuration achieved a Cohen’s kappa (κ) of 0.70 and 82% four-class accuracy. Structured reasoning prompting improved all Llama models but degraded all models with native reasoning capability (Qwen 3.5 and Gemma 4), revealing an architecture-dependent interaction. Thinking mode hurt performance on criteria-based prompts and never yielded the best configuration. Larger models did not consistently outperform smaller ones: Llama 3.1 405B offered no advantage over Llama 3.3 70B, and the Qwen 397B model underperformed the dense Qwen 27B. These findings demonstrate that newer model generations with improved general benchmarks and larger parameter counts do not guarantee better performance on specialized medical classification tasks and that prompt design must be matched to model architecture.