Introduction <p>The rapid uptake of large language models (LLMs) in surgery demands evidence of their reliability when guiding laparoscopic cholecystectomy (LC).</p> Methods <p>An analytical cross-sectional study (April–June 2025) compared five current LLMs (ChatGPT-o3, Claude-Sonnet-4, DeepSeek-V3.5, Gemini-2.5 Flash, and Grok-3) on 24 guideline-derived questions covering the pre-, intra-, and postoperative phases of LC. Four blinded hepatobiliary surgeons rated 120 answers with the eight-item modified DISCERN (mDISCERN, 8–40) and Global Quality Score (GQS, 1–5). Readability was quantified with FRES, FKGL, SMOG, Fog, CLI, and lexical density indices, and inter-rater agreement assessed by two-way ICC.</p> Results <p>Grok delivered the highest mean mDISCERN (36.3 ± 2.3) and GQS (4.76 ± 0.41), whereas Gemini scored lowest (29.0 ± 2.1; 3.58 ± 0.36). DeepSeek produced the most readable output (FRES&#xa0;≈&#xa0;30.6; FKGL&#xa0;≈&#xa0;12.1), while Claude generated the densest, least readable text (negative FRES; FKGL&#xa0;≈&#xa0;18.3). Quality correlated positively with word count and lexical density (<i>ρ</i>&#xa0;≈&#xa0;0.7) but not with syntactic complexity. Surgeon ratings showed good reliability (ICC(2,k) = 0.775; ICC(3,k) = 0.819).</p> Conclusions <p>LLM performance for LC varies markedly; even the best-performing model stops short of full reliability, reinforcing the need for procedure-specific validation before clinical deployment. This multidimensional audit provides a reproducible benchmark for selecting and fine-tuning surgical decision-support LLMs and highlights that terminological richness, rather than sentence complexity, underpins high-quality guidance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChatGPT and other large language models in laparoscopic cholecystectomy: a multidimensional audit of reliability, quality, and readability

  • Yusuf Yunus Korkmaz,
  • Oğuzhan Aydın,
  • Feyyaz Güngör,
  • İlyas Kudaş,
  • Talha Sarigoz,
  • Ozgur Bostanci

摘要

Introduction

The rapid uptake of large language models (LLMs) in surgery demands evidence of their reliability when guiding laparoscopic cholecystectomy (LC).

Methods

An analytical cross-sectional study (April–June 2025) compared five current LLMs (ChatGPT-o3, Claude-Sonnet-4, DeepSeek-V3.5, Gemini-2.5 Flash, and Grok-3) on 24 guideline-derived questions covering the pre-, intra-, and postoperative phases of LC. Four blinded hepatobiliary surgeons rated 120 answers with the eight-item modified DISCERN (mDISCERN, 8–40) and Global Quality Score (GQS, 1–5). Readability was quantified with FRES, FKGL, SMOG, Fog, CLI, and lexical density indices, and inter-rater agreement assessed by two-way ICC.

Results

Grok delivered the highest mean mDISCERN (36.3 ± 2.3) and GQS (4.76 ± 0.41), whereas Gemini scored lowest (29.0 ± 2.1; 3.58 ± 0.36). DeepSeek produced the most readable output (FRES ≈ 30.6; FKGL ≈ 12.1), while Claude generated the densest, least readable text (negative FRES; FKGL ≈ 18.3). Quality correlated positively with word count and lexical density (ρ ≈ 0.7) but not with syntactic complexity. Surgeon ratings showed good reliability (ICC(2,k) = 0.775; ICC(3,k) = 0.819).

Conclusions

LLM performance for LC varies markedly; even the best-performing model stops short of full reliability, reinforcing the need for procedure-specific validation before clinical deployment. This multidimensional audit provides a reproducible benchmark for selecting and fine-tuning surgical decision-support LLMs and highlights that terminological richness, rather than sentence complexity, underpins high-quality guidance.