Background <p>Multidisciplinary team (MDT) review is essential for individualized treatment planning in lung cancer, where decisions often require integration of staging, pathology, molecular findings, performance status, and multimodal treatment options. However, the effective implementation of MDT review remains constrained by limited specialist availability in resource-limited hospitals and by high patient volume, heavy physician workload, and tight clinical schedules in tertiary care centers. Large language models (LLMs) may support MDT workflows by synthesizing clinical information and generating preliminary treatment recommendations for clinician review. This study compared GPT-5 and Gemini-2.5-Pro with expert MDT recommendations in real-world lung cancer cases and evaluated how information completeness and input format influenced LLM-generated treatment recommendations.</p> Methods <p>We conducted a head-to-head evaluation of GPT-5 and Gemini-2.5-Pro for lung-cancer multidisciplinary team decision support. For each of 150 consecutively collected cases, structured vignettes were entered using three information-completeness levels in both natural-language and isomorphic JavaScript Object Notation (JSON) formats. Each model produced a single-sentence recommendation. Outputs were compared with the expert multidisciplinary team plan, which was treated as a clinical reference standard for concordance analysis rather than an objective guideline-only gold standard, using predefined plan-level categories. To better quantify model performance, we applied keyword extraction and semantic matching to normalize both multidisciplinary team recommendations and model outputs into ten treatment categories with subsequent manual review and correction.</p> Results <p>Stratifying by information level and format, Information Level III produced the best overall performance for Gemini-2.5-Pro under natural-language input, with F1 0.7603, accuracy 0.8647, precision 0.7031, and recall 0.8278. GPT-5 achieved its highest recall under Information Level III with JSON input (recall 0.9132), but this was accompanied by lower precision and F1. Under the same natural-language Information Level III, GPT-5 reached recall 0.8770 with F1 0.6826 and accuracy 0.7967.</p> Conclusions <p>Gemini-2.5-Pro show promise as adjuncts to MDT decision-making in lung cancer, offering high agreement with expert plans under controlled conditions. With appropriate clinician oversight, LLM-based tools may help expand access to standardized, high-quality recommendations where specialist availability is limited.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-assisted multidisciplinary decision-making in lung cancer: efficacy of Gemini-2.5-Pro and GPT-5 in a real-world MDT setting

  • Zichun Zhou,
  • Min Wang,
  • Jia Li,
  • Zhenchang Wang,
  • Han Lv

摘要

Background

Multidisciplinary team (MDT) review is essential for individualized treatment planning in lung cancer, where decisions often require integration of staging, pathology, molecular findings, performance status, and multimodal treatment options. However, the effective implementation of MDT review remains constrained by limited specialist availability in resource-limited hospitals and by high patient volume, heavy physician workload, and tight clinical schedules in tertiary care centers. Large language models (LLMs) may support MDT workflows by synthesizing clinical information and generating preliminary treatment recommendations for clinician review. This study compared GPT-5 and Gemini-2.5-Pro with expert MDT recommendations in real-world lung cancer cases and evaluated how information completeness and input format influenced LLM-generated treatment recommendations.

Methods

We conducted a head-to-head evaluation of GPT-5 and Gemini-2.5-Pro for lung-cancer multidisciplinary team decision support. For each of 150 consecutively collected cases, structured vignettes were entered using three information-completeness levels in both natural-language and isomorphic JavaScript Object Notation (JSON) formats. Each model produced a single-sentence recommendation. Outputs were compared with the expert multidisciplinary team plan, which was treated as a clinical reference standard for concordance analysis rather than an objective guideline-only gold standard, using predefined plan-level categories. To better quantify model performance, we applied keyword extraction and semantic matching to normalize both multidisciplinary team recommendations and model outputs into ten treatment categories with subsequent manual review and correction.

Results

Stratifying by information level and format, Information Level III produced the best overall performance for Gemini-2.5-Pro under natural-language input, with F1 0.7603, accuracy 0.8647, precision 0.7031, and recall 0.8278. GPT-5 achieved its highest recall under Information Level III with JSON input (recall 0.9132), but this was accompanied by lower precision and F1. Under the same natural-language Information Level III, GPT-5 reached recall 0.8770 with F1 0.6826 and accuracy 0.7967.

Conclusions

Gemini-2.5-Pro show promise as adjuncts to MDT decision-making in lung cancer, offering high agreement with expert plans under controlled conditions. With appropriate clinician oversight, LLM-based tools may help expand access to standardized, high-quality recommendations where specialist availability is limited.